8.4 KiB
Local AI Coding Environment — Reference & Lessons Learned
Consolidates
local_ai_setup.mdandlocal_llm.md. Those files are superseded by this one.
Hardware
| Component | Spec |
|---|---|
| GPU | NVIDIA RTX 5070 Ti, 12 GB VRAM |
| RAM | 32 GB system RAM |
| OS | Windows 11 |
| WSL | WSL2 Ubuntu |
Software Stack
| Tool | Purpose | Notes |
|---|---|---|
| Ollama | Local model server | Serves OpenAI-compatible API at localhost:11434/v1 |
| OpenCode | Primary AI coding agent | Installed via npm (opencode-ai) |
| Aider | Alternative AI coding agent | pip: aider-chat |
| Claude Code | Cloud AI coding agent | Optional; can point at Ollama |
| VS Code | Editor | With Continue extension for inline chat |
Architecture
VS Code
├── Continue Extension ──────────────┐
├── Claude Code CLI ────────────────┤
└── OpenCode / Aider ───────────────┤
▼
Ollama
│
▼
Qwen3-Coder 30B (primary)
All code stays local. Cloud models are optional fallback only.
Ollama — Available Models
| Model | Size | Use |
|---|---|---|
qwen3-coder:30b-fixed |
18 GB | Primary — OpenCode agent (fixed config, see below) |
qwen3-coder:30b |
18 GB | Original pull (unmodified) |
gemma4:latest |
9.6 GB | English-reliable alternative |
gemma4:12b |
7.6 GB | Lighter English-reliable option |
qwen2.5-coder:14b |
9.0 GB | Mid-weight coding model |
qwen2.5-coder:7b |
4.7 GB | Fast, lightweight coding |
qwen2.5-coder:1.5b |
986 MB | Minimal footprint |
qwen3:latest |
5.2 GB | General chat |
qwen3:1.7b |
1.4 GB | Fast general chat |
bge-m3:latest |
1.2 GB | Embeddings |
OpenCode Configuration
Config file: ~/.config/opencode/opencode.jsonc
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"ollama": {
"options": {
"baseURL": "http://localhost:11434/v1"
},
"models": {
"qwen3-coder:30b-fixed": { "name": "Qwen3 Coder 30B" },
"qwen2.5-coder:14b": { "name": "Qwen 2.5 Coder 14B" },
"qwen2.5-coder:7b": { "name": "Qwen 2.5 Coder 7B" },
"gemma4:latest": { "name": "Gemma 4 8B (latest)" },
"qwen3:latest": { "name": "Qwen3 8B" }
}
}
},
"model": "ollama/qwen3-coder:30b-fixed"
}
Active Model — qwen3-coder:30b-fixed Modelfile
FROM <local blob: sha256-9124a16eb1a7...>
TEMPLATE {{ .Prompt }}
PARAMETER repeat_penalty 1.05
PARAMETER stop <|im_start|>
PARAMETER stop <|im_end|>
PARAMETER stop <|endoftext|>
PARAMETER temperature 0.7
PARAMETER top_k 20
PARAMETER top_p 0.8
PARAMETER num_ctx 32768 ← fixed (was missing)
PARAMETER num_predict 8192 ← fixed (was missing)
Model specs: 30.5B params, MoE architecture (qwen3moe), Q4_K_M quantization, supports tools + completion.
Issue & Fix — Tool Call JSON Truncation
Error
llama-server returned invalid tool call arguments for "edit":
unexpected end of ... [retrying in 51s attempt N]
Root Cause
Two Ollama parameters were missing from the model's Modelfile:
| Parameter | Default when missing | Required for tool calls |
|---|---|---|
num_ctx |
2048 tokens | Minimum ~8192; 32768 recommended |
num_predict |
128 tokens | Minimum ~1024; 8192 recommended |
With only 128 output tokens allowed, the model began generating valid JSON for the edit tool call but ran out of budget before closing all braces, producing malformed JSON. The 2048-token context also filled up quickly during long sessions, compounding the problem.
Fix Applied
Rebuilt qwen3-coder:30b-fixed with the two parameters added:
# Export current Modelfile, append fixes, rebuild
ollama show qwen3-coder:30b-fixed --modelfile > Modelfile_fixed
# (add num_ctx 32768 and num_predict 8192)
ollama create qwen3-coder:30b-fixed -f Modelfile_fixed
No changes needed to the OpenCode config — same model name, same config file.
Verification
ollama show qwen3-coder:30b-fixed
# Should show: num_ctx 32768 and num_predict 8192 in Parameters section
Aider Configuration
Config file: .aider.conf.yml in project root
model: ollama/qwen3-coder:30b
openai-api-base: http://localhost:11434/v1
openai-api-key: ollama
edit-format: whole
map-tokens: 2048
no-auto-commits: true
system-prompt: "Always respond in English only. Never use Chinese or any other language."
| Option | Why |
|---|---|
edit-format: whole |
More reliable file writes for local models |
map-tokens: 2048 |
Limits repo map size for smaller context windows |
no-auto-commits: true |
Review changes before they hit git |
system-prompt |
Qwen3 defaults to Chinese; forces English |
Aider Quick Reference
| Command | What it does |
|---|---|
/add src/foo.py |
Add file to context |
/drop src/foo.py |
Remove file from context |
/diff |
See pending changes |
/undo |
Revert last change |
/run pytest |
Run command, share output with model |
/clear |
Clear chat history (keep files) |
/settings |
Show loaded config |
/model ollama/gemma4:12b |
Switch model on the fly |
Known Issues & Fixes
Tool call JSON truncated (OpenCode / llama-server)
Error: llama-server returned invalid tool call arguments for "edit": unexpected end of...
Cause: num_predict too low (default 128) cuts off JSON mid-generation.
Fix: See "Issue & Fix" section above — add num_ctx and num_predict to Modelfile.
Model responds in Chinese
Cause: Qwen3 is a Chinese model and defaults to Chinese.
Fix (Aider): system-prompt in .aider.conf.yml.
Fix (OpenCode): Add a system prompt in OpenCode's project config if supported, or use gemma4 models which default to English.
Aider writes code to chat but not to disk
Cause: Model didn't follow Aider's edit format.
Fix 1: /add target_file.py before the request — hints the target.
Fix 2: edit-format: whole in config (already included above).
Config changes not taking effect (Aider)
Fix: Exit (/exit) and restart. Config is read only at startup.
"File not found" / repo-map warnings for deleted files
Cause: File deleted from disk but still tracked in git index. Fix:
git rm --cached filename.py
# or from inside Aider:
/run git rm --cached filename.py
Model loops with filler text, never executes tool calls
Symptom: Model repeatedly says "I'll continue..." but never reads/edits files. Context counter in OpenCode bottom-left shows a number near your num_ctx limit.
Cause: Context window is full. No tokens left for tool call JSON, so the model outputs short filler and stalls.
Fix (immediate): Start a new OpenCode session. Break large tasks (e.g. "translate all files") into one file per session.
Fix (permanent): Increase num_ctx in the Modelfile (e.g. 65536), rebuild, restart Ollama. Watch RAM usage — each doubling of num_ctx roughly doubles KV cache memory.
Context fills up, model performance degrades
Cause: Long session accumulates tokens beyond num_ctx.
Fix (Aider): /clear to reset chat history without losing file context.
Fix (Ollama model): Increase num_ctx in Modelfile (current: 32768).
Performance Notes
- qwen3-coder:30b is a MoE model — fits in VRAM+RAM split on RTX 5070 Ti (12 GB VRAM).
num_gpu 99in Modelfile maximizes layers on VRAM; remainder spills to RAM.- Safe
num_ctxrange for this hardware: 8192–32768. Above 32768 risks RAM pressure. - Q4_K_M quantization gives best quality/speed/memory balance for this size class.
Environment Verification Checklist
ollama list # models appear
ollama show qwen3-coder:30b-fixed # num_ctx and num_predict present
curl http://localhost:11434/api/tags # Ollama serving
aider --version # Aider installed
opencode --version # OpenCode installed
nvidia-smi # GPU visible