Small Models. Full Control. No Cloud Required.

Timothy Gaul

โ€ข edge-ai, privacy, local-llms, web-assembly, machine-learning

The dominant AI narrative is all about scale - bigger models, bigger GPUs, bigger APIs.

But what if the future of AI isn't in the cloud - but on your device?

I've been diving into the world of 100M-1.7B parameter models like Phi-1.5 and TinyLLaMA. And what's most exciting isn't just how surprisingly capable they are, it's where they can run.

  • Directly in the browser using WebAssembly and WebGPU
  • Entirely on your phone, no internet connection required
  • Without sending a single token to OpenAI, Anthropic, or anyone else

This isn't theoretical. It's technically viable today and it completely flips the privacy model we've all accepted.

With quantisation and smart design, these compact models can:

  • ๐—ฆ๐˜‚๐—บ๐—บ๐—ฎ๐—ฟ๐—ถ๐˜€๐—ฒ ๐—ฎ๐—ป๐—ฑ ๐—ฐ๐—น๐—ฎ๐˜€๐˜€๐—ถ๐—ณ๐˜† your emails
  • ๐—ฆ๐˜‚๐—ด๐—ด๐—ฒ๐˜€๐˜ ๐—ฟ๐—ฒ๐—ฝ๐—น๐—ถ๐—ฒ๐˜€ ๐—ฎ๐—ป๐—ฑ ๐˜๐—ฎ๐—ด content
  • ๐—ฅ๐˜‚๐—ป ๐—ฅ๐—”๐—š ๐—ฝ๐—ถ๐—ฝ๐—ฒ๐—น๐—ถ๐—ป๐—ฒ๐˜€ on your encrypted local data
  • ๐—”๐—ป๐—ฑ ๐—ฑ๐—ผ ๐—ฎ๐—น๐—น ๐—ผ๐—ณ ๐˜๐—ต๐—ถ๐˜€ without ever touching the cloud

There's no prompt leakage. No API call logs. No vendor surveillance.

I'm building a privacy-first framework where the LLM, memory, context, and agent runtime all stay local - using technologies like WebLLM, IndexedDB, and lightweight JS-based orchestration. Think LangGraph, but offline and on-device.

We don't have to give up autonomy to get smart assistants.

If you're excited by the idea of edge-native, user-controlled AI, I'd love to connect.