Copilot CLI Gets Local Models: /model Discovers Your Ollama — but Telemetry Stays On, Offline Is Separate
On October 7, 2026, GitHub announced via Changelog: starting with CLI 1.0.94-0, the /model command discovers models in your local Ollama instance, listed alongside configured and cloud models. Discovery doesn't auto-enroll — each model needs manual confirmation — and models must support tool calling and streaming. GitHub also teased intelligent routing, and clarified: a local model neither enables offline mode nor disables telemetry.

On October 7, 2026, GitHub shipped a small but significant Changelog entry: starting with version 1.0.94-0, Copilot CLI can discover and use models from your local Ollama instance. Type /model in the CLI and it lists the models available in your running Ollama, right alongside your configured models and GitHub's cloud models.
The update is tiny — no new model, no new hotkey — yet it marks a philosophical turn for Copilot: for the first time, the coding assistant's "brain" doesn't have to be GitHub's cloud. It can be your own machine.
How It Works: Discovery ≠ Auto-Enrollment
The flow is deliberately cautious, in three steps: /model discovers models in your local Ollama → you review each model's provider and endpoint → you manually confirm "Add and use for this session" or "Add without switching." Discovery never auto-enrolls; every model needs your explicit confirmation, and once added it works in the current session without restarting the CLI.
Two hard prerequisites, stated plainly in the announcement: Ollama and the models must already be installed by you — this flow installs no runtime and downloads no models; and models must support tool calling and streaming, or they won't appear. Provider connection failures show up in the picker with an explanation, so you know what to fix.
The design extends the provider experience from the Copilot App (Settings → Model providers) into the CLI with the same logic: which model does the work is your call.
Three Honest Sentences Hidden in the Fine Print
The Changelog text is short, but three sentences deserve a close read:
First, "choosing a local model doesn't turn on offline mode or disable GitHub telemetry." Offline remains an explicit choice: COPILOT_OFFLINE=true. In other words, even when inference runs on your machine, prompts and code context may still travel to remote servers — a local model ≠ local processing of your code. Teams with compliance requirements should read that three times.
Second, GitHub teased "intelligent routing with local models." Details point to an announcement on the Microsoft Command Line blog. Intelligent routing means simple tasks automatically go to small local models while hard ones go to the cloud — if it lands well, it will be Copilot's first automatic scheduling across cost, speed, and quality.
Third, tool calling is a hard gate. Only models supporting tool calling make the list, which effectively draws a line: pure completion models need not apply; it has to be a proper agentic model. Plenty of Ollama-ecosystem models support tool calling, with uneven quality — run a real task before choosing, and don't judge by parameter count alone.
Why Now: Three Forces Converging
This update isn't random — three trends matured at once:
One: local models are genuinely useful now. In 2026, local models' usefulness on coding tasks is unrecognizable from two years ago — fixing a small bug, writing a test, explaining legacy code, a local 30B-class model handles it. Cloud giants keep architecture design and the truly nasty debugging; the division of labor forms naturally.
Two: cost pressure. Agentic coding burns multiples of the tokens completion-era coding did. Routing every high-frequency small task (formatting, comment fills, test-fix loops) through the cloud shows up visibly on the bill. Local models do this "manual labor" at near-zero marginal cost.
Three: data sovereignty. More teams now require "code never leaves the intranet," and local models are the only answer today — though, as noted above, Copilot's telemetry terms discount that answer; truly air-gapped scenarios still await the offline-mode evolution.
Practical Advice for Vibe Coders
1. Split traffic before replacing. Don't swap your primary model for a local one on day one. The suggested division: local models for high-frequency, low-risk tasks (test fixes, formatting, simple refactors, doc generation); cloud giants for architecture decisions and complex debugging. Once intelligent routing goes GA, this split may become automatic.
2. Make models pass the "tool calling" exam first. After pulling a model into Ollama, throw it a real task needing 3–5 consecutive tool calls (e.g., "add tests for this module and make them pass") and watch whether it drops the chain midway. Tool-calling stability matters far more than benchmark scores.
3. Read the telemetry terms before talking compliance. If your code is confidential, remember: picking a local model in /model ≠ data staying on your network. Read the CLI provider and offline-mode docs, confirm COPILOT_OFFLINE behavior and exactly what still gets sent, line by line, before bringing it into an intranet project.
4. Watch intelligent routing. It's the most interesting thing on Copilot's near-term roadmap — automatic "cheap local model leads, cloud giant backs up" scheduling, done well, could halve an indie developer's token bill.
Ollama-Side Prep: a Three-Step Readiness Check
Copilot CLI only "discovers" — Ollama and the models are on you. Three readiness checks:
1. Ollama is running. ollama serve starts the service; ollama list confirms models are pulled locally. The CLI discovers models "in a running Ollama instance" — service down, empty list.
2. The model supports tool calling. The hard gate. Verify in Ollama first with a quick run: have it call a tool (read a file, run a command) and watch whether it completes a multi-step chain. Pure completion models and legacy architectures are out.
3. Enough VRAM/RAM. Hardware sets the experience floor for local models: starved VRAM means brutal slowdowns, past a point worse than the cloud. Time a real task first — if a simple refactor takes 3 minutes locally versus 40 seconds in the cloud, the split makes no sense.
Walk It Through: From /model to Your First Local Task
The full flow looks like this: upgrade the CLI to 1.0.94-0; get ollama serve up with models in place; type /model in the CLI, find your local model, verify provider and endpoint, then pick "Add and use for this session"; throw it a low-risk trial task — e.g. "add unit tests for this function and make them pass" — no CLI restart needed.
Recommended trial ladder, three rungs: rung one is pure generation (tests, docs) — read-only or low risk; rung two is small edits (fix a known bug) — watch tool-calling stability; rung three is real requirements. If any rung breaks (dropped tool chains, heavy hallucination, unacceptable speed), fall back to the cloud model — adopting local models should be gradual, not a leap of faith.
Intelligent Routing Preview: Copilot's Next Move
The Changelog hides one more teaser: "intelligent routing with local models," details pointing at the Microsoft Command Line blog. Taken literally, it's Copilot's first attempt at automatic dispatch by task difficulty: simple tasks answered instantly by small local models, hard ones routed to cloud giants.
If it lands well, it kills vibe coders' biggest hidden cost: decision fatigue. Today every task starts with "is this worth the big model?" — automated routing erases that decision and optimizes the bill by itself. It's the industry's shared direction too (everyone's routing features and small-model price cuts follow the same logic): the second half of 2026 competes on scheduling, not on single models.
But don't architect around it until the official details drop. A teaser is a teaser — wait for GA. The vibe-project iron rule: never chase previews, only GAs.
Versus the Copilot App's Provider Experience
This CLI update essentially backports the provider experience already in the Copilot App (Settings → Model providers) to the terminal. Same logic both sides: "which model does the work is your call." The difference is scenario: the App serves interactive IDE coding; the CLI serves agentic terminal workflows — exactly where local models shine (batched small tasks, scripted invocations).
One signal worth noting: GitHub is handing "model choice" back from the product to the user. From the cloud-model bundle, to the App's third-party providers, to the CLI's local Ollama — the trajectory is clear: Copilot is shifting from "selling models" to "selling scheduling." That's good for users: the more models compete, the more the scheduling layer is worth — and you're standing on the cashier's side.
Do the Math: How Much Do Local Models Actually Save?
Rough numbers. Say your agent runs 50 small tasks a day (tests, lint fixes, docs) at ~30k tokens each. A mid-tier cloud model at roughly $3 per million tokens costs $4.50 a day — about $100 a month over 22 workdays. If a local model handles 70% of those tasks, the monthly bill drops to $30 — the $70 saved buys a decent hard drive.
But the other side of the ledger is hardware and power: a machine that comfortably runs 30B-class models starts at several thousand dollars in GPU/RAM, plus always-on electricity. The verdict: if you already have idle compute (a dev machine, a NAS, an old GPU), local models are pure profit; if you're buying cards to save tokens, the payback period is measured in years — try existing hardware first; don't spend money to save money.
And one ledger entry that's hard to quantify: latency. Local models skip the network round trip; small tasks often finish end-to-end faster than cloud calls — and indie developers all know what "instant reply" is worth to flow state.
One last angle: this shipped in the Changelog's "Improvement" section on October 7, not as a standalone blog fanfare — inside GitHub, local-model support is evidently "a matter of course," not a strategic launch. When big companies start announcing features like this in an Improvement tone, it means local models are genuinely usable, not demo-ware.
One-line summary: Copilot CLI's local-model support marks the coding assistant's move from "one cloud to rule them all" to "hybrid scheduling." The brain can live locally — but read the terms: where the model runs and where your data goes are two different things.
Sources
Related articles

On October 7, 2026, Google Developers launched the Developer Knowledge API ecosystem: official Google Cloud, Firebase, and Android docs as a programmatic source of truth, with a gcloud CLI surface, an official Agent Skill (one-line install), an MCP server, and multi-language client libraries. Why 'docs as APIs' uproots vibe coding's classic failure of models misremembering APIs.

On October 3, engineer Kevin Liao published a polemic that hit the HN front page: agent memory plugins are a lottery over RAG snippets; what agents need is a documentation workspace. The essay's diagnosis, its open-source Operator Memory plugin, the two strongest objections, and the minimal practice you can start tonight.

On October 8, 2026, Google Cloud launched the Gemini agent at Gemini at Work 2026: a universal agent for work that takes objectives, plans by itself, auto-selects between Gemini and Claude models per task, and introduces 'coworker agents' with their own email, calendar, and directory seat. Four judgments on why the second half of the agent race is about 'agents that feel like colleagues.'