01 article

Antigravity's cloud planner spends 95 tokens, local Gemma does the rest

Google's Antigravity SDK now runs agents fully offline on local Gemma 4 26B, but the real story is the hybrid split: a cloud planner spends 95 tokens, and local models handle 97.2% of the rest.

Antigravity's cloud planner spends 95 tokens, local Gemma does the rest

Antigravity's cloud planner spends 95 tokens, local Gemma does the rest

Google quietly moved the needle on something most local-model posts skip. On September 23 the Antigravity SDK, the same agentic engine behind Google's Antigravity IDE, learned to run entirely on your own hardware. You can point it at a local Gemma 4 26B and it will plan, edit, test, and ship code without a single request leaving your machine. The offline part is the easy story. The part that got me is how they split the work between a cloud model and a local one.

I've watched "run an agent locally" get marketed as a privacy feature, a cost feature, an offline feature, and a toy, usually in that order. Google's post does all four, but the hybrid orchestration demo is where the idea actually lands.

What shipped

The SDK now takes a local model through a new LiteRTAgentConfig, backed by Google AI Edge's LiteRT runtime. The reference model is Gemma 4 26B (the A4B variant), pulled in from Hugging Face with a one-line litert-lm import. Google recommends at least 24GB of VRAM or unified memory, so this is a desktop-Mac or beefy-GPU story, not a laptop-and-coffee story.

pip install google-antigravity litert-lm

litert-lm import \
  --from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
  gemma-4-26B-A4B-it-gpu.litertlm \
  gemma4-26b

Then the agent is just a config and a loop:

config = LiteRTAgentConfig(model_path=MODEL_PATH).lightweight()
async with Agent(config) as agent:
    response = await agent.chat("What files are in the current directory?")
    async for token in response:
        print(token, end="", flush=True)

The less obvious win is LocalOpenAIAgentConfig. Any OpenAI-compatible server, Ollama, LM Studio, or vLLM, plugs in as the model and your orchestration, tools, and workflows stay identical. You swap the model, not the agent.

The 95-token architect

Here's the pattern I want to steal. Google calls it the Architect-Builder split. A cloud model, Gemini 3.8 Flash, acts as the planner and conductor. A swarm of local Gemma 4 26B instances does the actual work on the GPU. In their demo the task is to audit and patch three vulnerable Python modules: auth.py, billing.py, and database.py.

Two numbers do the talking. The cloud architect spent 95 tokens total. It planned the strategy and decomposed the work from filenames and task descriptions alone, and no source code ever left the machine. Then the local models took over for an adversarial gauntlet: reproduce the vulnerability, write a candidate fix, critique the patch, and validate it against the regression suite. Across the whole run, 97.2% of all tokens, 3,322 of them, executed locally and offline, and the patches came back green.

The security angle is why the gauntlet is shaped the way it is. A local model that both writes and reviews its own patch sounds circular, so the loop forces reproduction first. A fake fix that does not actually close the vulnerability gets caught by the test suite before it ever counts as a pass.

I keep coming back to that 95. It is a concrete expression of a split that most teams hand-wave about. The expensive frontier model handles the part that actually needs scale, the planning, and the commodity local model absorbs the part that is mostly grind, the reproduction, the drafting, the critique. You pay cloud money for judgment and local power for labor.

When local agents earn their keep

The honest reasons to do this are the same four Google lists, but weighted differently than the marketing implies. Privacy is the one that is real for enterprises, keeping proprietary code out of a third-party API is a compliance decision, not a preference. Cost matters when the loop is churning tokens on boilerplate that a small model can handle. Offline matters for the odd air-gapped box or flaky field connection. The hybrid framing is the durable one: route the cheap steps locally and escalate only what genuinely needs the big model.

In practice that split looks like a routing table. Summarize the diff locally, let the cloud model decide which parts are risky, then push the risky branches back through the local model against your own tests. The frontier model never sees the full codebase, and the small model never has to make a judgment call it is bad at.

The limits are worth stating plainly. The 24GB floor keeps this off most laptops. The token accounting is Google's, from a recorded demo run, so treat 97.2% as directional until you can reproduce it on your own modules and test suite. And the A4B build is a specific model choice, not a guarantee that any local model will hold an adversarial audit loop together. A 26B local model critiquing its own patches is a demo of a pattern, and it will wobble on harder codebases.

How to try it

Google ships an example project with a built-in three-file gauntlet. Run it as-is to see the loop, or point it at your own Python modules and tests, which is where it stops being a demo and starts being a tool. You do not need to replicate their exact stack; the point is the harness, and the harness now has a first-class local path.

The CLI-utility example in the post is a nice sanity check. One prompt, and the local agent wrote a live psutil and rich dashboard, generated the requirements file, and tested it, all offline. If a 26B local model can do that, it can probably do a lot of the chores your team currently burns API tokens on.

What this really signals is that local inference stopped being the fallback tier. It is becoming a routing decision inside agent orchestration, with the cloud model as the brain and local models as the hands. I would watch whether the Architect-Builder split shows up in non-Google stacks next, because 95 cloud tokens for a full plan and local horsepower for the rest is an economics argument that does not need a Google logo to be true.

Comments