Files
2026-06-10 19:45:12 +08:00

8.4 KiB
Raw Permalink Blame History

Local AI Coding Environment — Reference & Lessons Learned

Consolidates local_ai_setup.md and local_llm.md. Those files are superseded by this one.


Hardware

Component Spec
GPU NVIDIA RTX 5070 Ti, 12 GB VRAM
RAM 32 GB system RAM
OS Windows 11
WSL WSL2 Ubuntu

Software Stack

Tool Purpose Notes
Ollama Local model server Serves OpenAI-compatible API at localhost:11434/v1
OpenCode Primary AI coding agent Installed via npm (opencode-ai)
Aider Alternative AI coding agent pip: aider-chat
Claude Code Cloud AI coding agent Optional; can point at Ollama
VS Code Editor With Continue extension for inline chat

Architecture

VS Code
  ├── Continue Extension  ──────────────┐
  ├── Claude Code CLI  ────────────────┤
  └── OpenCode / Aider  ───────────────┤
                                        ▼
                                     Ollama
                                        │
                                        ▼
                               Qwen3-Coder 30B (primary)

All code stays local. Cloud models are optional fallback only.


Ollama — Available Models

Model Size Use
qwen3-coder:30b-fixed 18 GB Primary — OpenCode agent (fixed config, see below)
qwen3-coder:30b 18 GB Original pull (unmodified)
gemma4:latest 9.6 GB English-reliable alternative
gemma4:12b 7.6 GB Lighter English-reliable option
qwen2.5-coder:14b 9.0 GB Mid-weight coding model
qwen2.5-coder:7b 4.7 GB Fast, lightweight coding
qwen2.5-coder:1.5b 986 MB Minimal footprint
qwen3:latest 5.2 GB General chat
qwen3:1.7b 1.4 GB Fast general chat
bge-m3:latest 1.2 GB Embeddings

OpenCode Configuration

Config file: ~/.config/opencode/opencode.jsonc

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ollama": {
      "options": {
        "baseURL": "http://localhost:11434/v1"
      },
      "models": {
        "qwen3-coder:30b-fixed": { "name": "Qwen3 Coder 30B" },
        "qwen2.5-coder:14b":     { "name": "Qwen 2.5 Coder 14B" },
        "qwen2.5-coder:7b":      { "name": "Qwen 2.5 Coder 7B" },
        "gemma4:latest":         { "name": "Gemma 4 8B (latest)" },
        "qwen3:latest":          { "name": "Qwen3 8B" }
      }
    }
  },
  "model": "ollama/qwen3-coder:30b-fixed"
}

Active Model — qwen3-coder:30b-fixed Modelfile

FROM <local blob: sha256-9124a16eb1a7...>
TEMPLATE {{ .Prompt }}
PARAMETER repeat_penalty  1.05
PARAMETER stop            <|im_start|>
PARAMETER stop            <|im_end|>
PARAMETER stop            <|endoftext|>
PARAMETER temperature     0.7
PARAMETER top_k           20
PARAMETER top_p           0.8
PARAMETER num_ctx         32768      ← fixed (was missing)
PARAMETER num_predict     8192       ← fixed (was missing)

Model specs: 30.5B params, MoE architecture (qwen3moe), Q4_K_M quantization, supports tools + completion.


Issue & Fix — Tool Call JSON Truncation

Error

llama-server returned invalid tool call arguments for "edit":
unexpected end of ... [retrying in 51s attempt N]

Root Cause

Two Ollama parameters were missing from the model's Modelfile:

Parameter Default when missing Required for tool calls
num_ctx 2048 tokens Minimum ~8192; 32768 recommended
num_predict 128 tokens Minimum ~1024; 8192 recommended

With only 128 output tokens allowed, the model began generating valid JSON for the edit tool call but ran out of budget before closing all braces, producing malformed JSON. The 2048-token context also filled up quickly during long sessions, compounding the problem.

Fix Applied

Rebuilt qwen3-coder:30b-fixed with the two parameters added:

# Export current Modelfile, append fixes, rebuild
ollama show qwen3-coder:30b-fixed --modelfile > Modelfile_fixed
# (add num_ctx 32768 and num_predict 8192)
ollama create qwen3-coder:30b-fixed -f Modelfile_fixed

No changes needed to the OpenCode config — same model name, same config file.

Verification

ollama show qwen3-coder:30b-fixed
# Should show: num_ctx 32768 and num_predict 8192 in Parameters section

Aider Configuration

Config file: .aider.conf.yml in project root

model: ollama/qwen3-coder:30b
openai-api-base: http://localhost:11434/v1
openai-api-key: ollama
edit-format: whole
map-tokens: 2048
no-auto-commits: true
system-prompt: "Always respond in English only. Never use Chinese or any other language."
Option Why
edit-format: whole More reliable file writes for local models
map-tokens: 2048 Limits repo map size for smaller context windows
no-auto-commits: true Review changes before they hit git
system-prompt Qwen3 defaults to Chinese; forces English

Aider Quick Reference

Command What it does
/add src/foo.py Add file to context
/drop src/foo.py Remove file from context
/diff See pending changes
/undo Revert last change
/run pytest Run command, share output with model
/clear Clear chat history (keep files)
/settings Show loaded config
/model ollama/gemma4:12b Switch model on the fly

Known Issues & Fixes

Tool call JSON truncated (OpenCode / llama-server)

Error: llama-server returned invalid tool call arguments for "edit": unexpected end of... Cause: num_predict too low (default 128) cuts off JSON mid-generation. Fix: See "Issue & Fix" section above — add num_ctx and num_predict to Modelfile.

Model responds in Chinese

Cause: Qwen3 is a Chinese model and defaults to Chinese. Fix (Aider): system-prompt in .aider.conf.yml. Fix (OpenCode): Add a system prompt in OpenCode's project config if supported, or use gemma4 models which default to English.

Aider writes code to chat but not to disk

Cause: Model didn't follow Aider's edit format. Fix 1: /add target_file.py before the request — hints the target. Fix 2: edit-format: whole in config (already included above).

Config changes not taking effect (Aider)

Fix: Exit (/exit) and restart. Config is read only at startup.

"File not found" / repo-map warnings for deleted files

Cause: File deleted from disk but still tracked in git index. Fix:

git rm --cached filename.py
# or from inside Aider:
/run git rm --cached filename.py

Model loops with filler text, never executes tool calls

Symptom: Model repeatedly says "I'll continue..." but never reads/edits files. Context counter in OpenCode bottom-left shows a number near your num_ctx limit. Cause: Context window is full. No tokens left for tool call JSON, so the model outputs short filler and stalls. Fix (immediate): Start a new OpenCode session. Break large tasks (e.g. "translate all files") into one file per session. Fix (permanent): Increase num_ctx in the Modelfile (e.g. 65536), rebuild, restart Ollama. Watch RAM usage — each doubling of num_ctx roughly doubles KV cache memory.

Context fills up, model performance degrades

Cause: Long session accumulates tokens beyond num_ctx. Fix (Aider): /clear to reset chat history without losing file context. Fix (Ollama model): Increase num_ctx in Modelfile (current: 32768).


Performance Notes

  • qwen3-coder:30b is a MoE model — fits in VRAM+RAM split on RTX 5070 Ti (12 GB VRAM).
  • num_gpu 99 in Modelfile maximizes layers on VRAM; remainder spills to RAM.
  • Safe num_ctx range for this hardware: 819232768. Above 32768 risks RAM pressure.
  • Q4_K_M quantization gives best quality/speed/memory balance for this size class.

Environment Verification Checklist

ollama list                          # models appear
ollama show qwen3-coder:30b-fixed    # num_ctx and num_predict present
curl http://localhost:11434/api/tags # Ollama serving
aider --version                      # Aider installed
opencode --version                   # OpenCode installed
nvidia-smi                           # GPU visible