Arun's blog

Notes on Local AI - Why I went local AI

Table of Contents

  1. The Engineer's Muscle: Embracing the Hard Problems
  2. The Craving for Extreme Personalization
  3. A Brief History of Personalization
  4. The Problem with "General Purpose" Software
  5. My Prototype: Hand-coding the "Artisan" Life
  6. The Tipping Point: When Cloud AI Breaks Down
  7. The Physics of Local vs. Cloud AI
  8. Local Models have become powerful
  9. Privacy and Latency
  10. Conclusion

The Engineer's Muscle: Embracing the Hard Problems

As a software engineer, I have spent several years handwriting code, thinking through complex hard problems, facing ambiguity, and breaking down complex problems into smaller executable chunks. Recently, I pivoted to ML pre-training and hardware-specific kernel programming. I was yet again faced with a complex problem which was squarely out of my comfort zone. This forced me to learn how to break down entirely new types of technical challenges.

As an engineer, I have always had control over the hardware and software I build. But when I step out of my job into my personal life and use consumer apps, I am suddenly at the mercy of rigid, "dumb" consumer apps.

The Craving for Extreme Personalization

The next disruption I wanted was in my personal life. There are a lot of things engineered around us with suboptimal personalized use cases for me. A general-purpose app from a marketplace solves the problem in general for most people, but my to-do app is not going to tell me exactly what I want to know for a given moment in time. I know what I want to know, but the app is dumb. This is not really just an app problem, it's a personalization problem. Assume you have one app per customer.

Is this really a problem? Do you really want personalization at an app level?

A Brief History of Personalization

To think about this, let's take a step back and look at the journey of personalization across all industries.

  1. The Artisan Era (Pre-Industrial Revolution)
    • Personalization wasn't a feature; it was just how things worked. Tailors made clothing specific to your body, the supermarket owner (which was just a local market, really) knew what goods your family needed, and the local cobbler and barber knew what your shoe needs were and the haircut styles you preferred.
    • It was highly localized to a small community and impossible to scale, and expensive, too.
  2. The Mass-Production Era (Late 19th to Mid-20th centuries)
    • Max out scale, minimize personalization (because it doesn't scale).
    • Efficiency, uniformity, and affordability.
    • As a consumer, you trade personalized experiences for lower cost and wider availability.
  3. The Mass Personalization Era (Late 20th century)
    • As competition increased, brands knew they had to differentiate.
    • Basic demographics like age, gender, and zip code entered the game. Voila! CRM is born. You are no longer an individual, but a bucket (e.g., college student, new dad).
    • Slightly more relevant, but still generic.
  4. The E-Commerce Era (2000 - 2010)
    • New tech called the internet now allowed companies to track clicks, purchase history, and search history.
    • An abandoned cart now sends you an email reminding you to complete your checkout.
    • But this is still heavily reliant on historical data and a lot of rigid "IFTTT" (If This, Then That) logic.
    • If you have ever felt creeped out by an ad, it was due to this. You get ads for shoes even though you already bought them.
  5. Hyper-Max-Personalization (Today)
    1. Predictive 1-to-1 experiences at scale.
    2. Spotify tells me what I will like, Netflix tells me what I will like watching, and Starbucks tells me what I need to drink based on the time of my day or the weather.

Across all industries, we are just going back to the Artisan Era to recreate that intimacy at scale.

The Problem with "General Purpose" Software

Hopefully, by now, I have convinced you that "General Purpose" IS the problem. If not, let me rehash it a little bit to make my point more solid.

  1. Software that resonates with your mind:
    • It's an app where you don't need to learn the interface. I like to plan my week by grouping it as 'Deep work', 'Errands', and 'Chores', and I never want to do chores on Fridays. The software instantly generates the UI, database, and logic for your specific model. If your workflow actually changes a month later (which it will, trust me), the software reshapes itself.
  2. Contextual benefits:
    • You don't have to know what else is going on in your life. You are not going to get a "It's time to stand up now" notification when you are almost asleep. Instead, you'll get: "You have 30 minutes before you have to go grocery shopping; here are 2 emails you can reply to now before that."
  3. Reduction in Cognitive load:
    • Right now you are the CEO of your own life, and apps are just how you run your business. Personalized software shifts the load from human to machine. You no longer manage the tool; the tool manages the complexity of your life.

So the disruption isn't really about finding the right app. It's about the death of the idea of a perfect app, replaced by fluid software that evolves with you and is as unique as your thoughts.

My Prototype: Hand-coding the "Artisan" Life

Let's look at my problem statement now. I was craving (borderline desperate) for personalization at a software level. This was a recent realization (say 2 years ago). I started coding my own apps and self-hosting them on local servers. This was before the ChatGPT boom. I hand-wrote code because no one could write it for me.

But wait, this doesn't scale. Let alone for normal people, even for me this doesn't scale. I can't write software long-term that makes it fluid for my use cases as I keep changing my mind.

Then the coding assistants became so good, and they eventually started to code better and faster than me. So I just ended up building a lot of apps personalized for myself and locally hosted them.

It has been a really fulfilling experience so far, especially when I am able to just add features that I need in under a day and have it already live and being used.

The Tipping Point: When Cloud AI Breaks Down

Note: I highly encourage you to stop reading here if you disagree with the premise above that personalized software is a structural necessity in the future. Otherwise, you will always view Local AI as just a hobbyist alternative to rigid apps.

How do you use AI? Do you use it as a search box you query occasionally, or is it the operating system of your personal life?

When you go from one end of that spectrum to the other, a lot of the fundamental physics regarding how we compute and pay for software changes.

The Physics of Local vs. Cloud AI

Your AI is always ON - Can you pay for that?

Now that you live in a world of generative, hyper-personalized UI, your agent isn't just generating text when you ask a question. It's always ON in the background: looking at your location, time, email, and calendar, updating your tasks on the fly, and updating your UI on the fly.

Can you really go Cloud on this? Say you rely on a cloud provider like OpenAI or Anthropic. Every single state change, background task, and UI generation requires an API call. You pay per token because your app is "constantly thinking" about what you need next. If you have set this up correctly, you will instantly burn through your CapEx for AI.

What if you go local? Zero marginal cost for inference. You have an upfront CapEx of buying Apple Silicon or a PC with a sick GPU. Now the cost is just electricity. Inference is constant and default like how you breathe. You forget you are doing it, and it just works foundationally.

The Efficiency Trap - but only the economics, not the brain rot

The paradox with cloud-based AI is: the more efficient and integrated you are, the more it will cost you. 🫠

This also creates a nice segue to the next item.

The Cost of Changing Your Mind

Personal life is dynamic. Humans make mistakes and change their minds. I have thoughts today that I know I will disagree with in the past.

Local Models have become powerful

A year ago, I would have strongly disagreed with going local. I bought my hardware because I wanted to play around and learn. The goal was not local AI.

And finally...

Privacy and Latency

If it's going to truly manage your personal life, it needs access to intimate data: health, finance, private journals, and real-time location. I would rather not upload myself to a cloud.

And the latency of a roundtrip to a cloud breaks the illusion of fluid UI. Local feels significantly less sluggish than cloud.

Conclusion

All of the above hopefully explains why I chose to go to Local AI. I also believe this helps you think about your relationship with AI and how you want that to be in the future.

P.S. From some quality feedback I got from my previous post, I am moving the fun AI slop section below as "Nerd area" for folks who are interested in the actual math of my decision-making.

To conclude, if you want true personalization, you need an always-on OS. It needs to be local, otherwise you are just going to settle for a lifetime of subscriptions (rented intelligence) that actively penalize you financially the more productive you become.


WARNING !!! ::: AI slops below :::

Nerd Area: The Local AI Tipping Point Equation

For the engineers who refuse to make a life choice without defining a mathematical threshold, here is the formula to calculate your exact tipping point for abandoning the cloud.

We define Δlocal as the Annualized Local Advantage. If Δlocal>0, it is statistically irrational for you not to buy a massive GPU or an Apple Silicon Mac today.

Δlocal=[365·k=1Ntasks(Nk·Ik·Ck)]Cloud OpEx(LcDy+k=1NtasksEk)Local Hardware CapEx + Power+(Ps·Vd)Privacy Premium+((LcloudLlocal)·It)LatencyGain+(Fa·Sc)Alignment Friction+Wjoy

The 16 Variables Explained:

The Economics (Cloud vs. Metal):

  1. Nk = The daily frequency of a specific type of task k (both explicit queries and "Always-ON" background events).
  2. Ik = Multimodal Intensity Multiplier. A weighted variable for task complexity. Parsing text = 1x. Synthesizing a personalized Text-to-Speech (TTS) news podcast = 50x. Generating a bespoke UI rendering or video = 500x.
  3. Ck = The blended cloud API cost for that specific task (input/output tokens, image generation credits, or audio seconds).
  4. Lc = Capital Expenditure of local hardware (The upfront cost of that sweet M3 Max or RTX 4090 rig).
  5. Dy = Depreciation lifecycle of your hardware in years (usually 3 to 5).
  6. Ek = Local electricity cost specifically mapped to the GPU/NPU utilization of task k.

The Intangibles (Privacy & Latency): 7. Ps = Privacy sensitivity coefficient (Scale of 0 to 1, where 1 means the AI is reading your real-time GPS and banking data). 8. Vd = The monetary value you place on your personal data not being part of a cloud provider's next pre-training run. 9. Lcloud = Average roundtrip latency of cloud API calls (ms). 10. Llocal = Average Time-To-First-Token (TTFT) or render time of your local model (ms). 11. It = Your personal intolerance threshold for sluggish UI (measured in dollars per millisecond of annoyance).

The "Changing Your Mind" Penalty: 12. Fa = Cloud alignment friction (the percentage of time the cloud API refuses a benign personal task because it tripped a generic safety filter). 13. Sc = The sunk cost (in dollars of your hourly rate) required to rethink and redesign your workflow every time a cloud provider deprecates your favorite model. 14. Wjoy = The pure, unadulterated psychological joy of owning your own model weights (Priceless, but for the sake of math, let's value it at $1,000/year).

The Conclusion of the Math: As long as the Multimodal Intensity (Ik) of your background tasks grows, which it must as your personalized OS starts reading to you, generating custom UIs, and rendering video, the Cloud OpEx scales exponentially. Therefore, Δlocal is guaranteed to be wildly positive. Go buy the hardware.

tbh, the above pretty much was N shotted by an AI, pretty impressive. I had to go back and forth a few times to get all my variables right, but its all there.