vision: real camera frames to the model (image_url), auto-detected for multimodal ollama models; --vision/--no-vision flags; verified gemma4:12b sees frames and completed the beacon mission

This commit is contained in:
opencode
2026-08-08 19:38:21 +03:00
parent 6a509aacc6
commit ef02e5f9ea
6 changed files with 5747 additions and 6 deletions
+14
View File
@@ -101,6 +101,20 @@ works: `python -m testbed.chat --base-url https://api.openai.com/v1
--model gpt-4o-mini --api-key $OPENAI_API_KEY`. `--auto-steps N` controls how
many tool steps the model may chain per request (0 = one action per turn).
### Real vision
By default the model gets frames as a color-grid digest (works for any
text-only model). If the model can actually see images (gemma3/gemma4,
qwen2.5-vl, gpt-4o-mini, ...), pass the frames as real images:
- auto-detected for local ollama (`/api/show` capabilities) — nothing to do
- otherwise force it: `--vision` (or `--no-vision` to disable)
With vision ON, `vision` tool results attach the actual camera frame as an
image to the conversation — the model sees the crate, the wall, the glowing
beacon, and orients itself. Verified locally: gemma4:12b completed a beacon
mission on its first attempt using look_at/move navigation.
## Real-time mode
Watch the capsule drive live in your browser: