Fix cross-process recall: MEMB trailer + raw eval/sample in chat()
Two bugs were blocking memba's main promise (load .memb in a fresh process → model continues with full recalled context): 1. llama_state_set_data() restores the C-level KV-cache + SSM hidden state, but llama-cpp-python's Python wrapper still reports n_tokens=0. The next eval() then decodes new tokens at offset 0 and overwrites the loaded state. Fix: extend MEMB format with an optional 12-byte trailer appended after the CRC32. It carries the wrapper's n_tokens. The C library reads up to CRC and ignores anything past it, so files stay backward-compatible with libmemba; only the Python loader uses it. 2. Llama.__call__ / create_chat_completion / generate all re-tokenise the prompt on every call and clear the KV-cache when the new tokens don't prefix-match input_ids. That destroys any state we just loaded. Fix: rewrite Session.chat() to use raw tokenize → eval → sample. eval() appends tokens to the live state without resetting, and we handle stop-token detection ourselves. Verified end-to-end on Nemotron-3-Nano-4B (hybrid 21x Mamba-2 + 4x attention) — see experiments/README.md for the full findings log. diag_session_nemotron.py, mood_batch_poc.py, mood_stream_poc.py and recall_poc.py now all pass their cross-process tests; Falcon-Mamba still fails because the trained model itself can't do cross-turn recall — that was the original misdiagnosis. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
+44
-19
@@ -1,32 +1,57 @@
|
||||
# experiments/
|
||||
|
||||
Throwaway scripts used to probe capabilities of SSM models with memba.
|
||||
Throwaway scripts used to probe SSM/hybrid model capabilities with memba.
|
||||
Not part of the library API — kept in the repo as reference and reproducible
|
||||
evidence for product decisions.
|
||||
|
||||
Each script is self-contained and prints what it finds; read the source for
|
||||
the test's claim and run it yourself if you want to verify on different
|
||||
models or hardware.
|
||||
Each script is self-contained. Read the source for what it claims to test
|
||||
and run it yourself if you want to verify on different models or hardware.
|
||||
|
||||
## Scripts
|
||||
|
||||
| File | What it measures |
|
||||
|------|------------------|
|
||||
| `recall_poc.py` | Can a memba state, built from N days of git activity, answer "what did I work on last month" in a fresh process? |
|
||||
| `mood_poc.py` | Batch sentiment-trajectory test (single prompt with full chat log). Compares Falcon-Mamba vs a Transformer. |
|
||||
| `mood_stream_poc.py` | The same trajectory but fed turn-by-turn through `Session.chat()`, then queried cross-process. |
|
||||
| `mood_batch_poc.py` | Batch ingest in build process, save, then query in a fresh process. |
|
||||
| `diag_saveload.py` | Minimal diagnostic: tell the model one fact, ask it back before save, after save, after cross-process load. |
|
||||
| `recall_poc.py` | git log → state → cross-process "what did I work on last month" |
|
||||
| `mood_poc.py` | Batch sentiment trajectory in one prompt — Falcon-Mamba vs Gemma |
|
||||
| `mood_batch_poc.py` | Batch ingest, save, query in fresh process |
|
||||
| `mood_stream_poc.py` | 15 separate observation turns, save, query in fresh process |
|
||||
| `diag_saveload.py` | Minimal hamster recall test on Falcon-Mamba |
|
||||
| `diag_nemotron.py` | Same hamster test on Nemotron 4B hybrid |
|
||||
| `diag_nemotron2.py` | Cross-process recall via raw `generate()` (bypass `create_chat_completion`) |
|
||||
| `diag_nemotron3.py` | Same with llama-cpp-python's native `save_state()`/`load_state()` |
|
||||
| `diag_session_nemotron.py` | Full hamster recall through the rewritten `Session.chat()` |
|
||||
|
||||
## Headline finding (2026-05-16, Falcon-Mamba-7B-Instruct Q4_K_M)
|
||||
## Findings log
|
||||
|
||||
- **Batch single-prompt analysis** (all input + question in one call): works
|
||||
for both sentiment and recall.
|
||||
- **Multi-turn fact recall** (ingest in turn 1, ask in turn 2): fails even
|
||||
in the *same process*. The model does not preserve specific facts in its
|
||||
hidden state across conversation turns.
|
||||
- **Save/load roundtrip**: byte-identical, no information loss attributable
|
||||
to memba's file format. The persistence layer works correctly; the
|
||||
trained model just doesn't use the state for cross-turn recall.
|
||||
### 2026-05-16 (initial, Falcon-Mamba-7B-Instruct Q4_K_M)
|
||||
|
||||
See the script outputs (or rerun) for the raw evidence.
|
||||
- Batch single-prompt analysis works for both sentiment and recall.
|
||||
- Multi-turn fact recall failed even in-process. *Misdiagnosed at first as
|
||||
a model capability issue.*
|
||||
|
||||
### 2026-05-16 (revised, after Nemotron-3-Nano-4B Q4_K_M test)
|
||||
|
||||
- Real root cause: `Llama.create_chat_completion()` and `Llama.generate()`
|
||||
retokenise the entire prompt each call and reset KV-cache when the new
|
||||
tokens don't prefix-match `input_ids`. Loaded memba state was being wiped.
|
||||
- Secondary issue: `llama_state_set_data()` restores the C-level cache but
|
||||
llama-cpp-python's wrapper still reports `n_tokens=0`, so the next
|
||||
`eval()` writes new tokens at offset 0 and overwrites the loaded state.
|
||||
- Two fixes in memba:
|
||||
1. MEMB file now carries a Python-only trailer with `n_tokens` so the
|
||||
wrapper position is restored after `load_state` (12 extra bytes,
|
||||
backward-compatible — the C library ignores anything after the CRC).
|
||||
2. `Session.chat()` rewritten to use raw `tokenize → eval → sample`
|
||||
instead of `Llama.__call__`, avoiding the prefix-matching reset.
|
||||
- Results on Nemotron-3-Nano-4B (hybrid: 21 Mamba-2 + 4 attention layers)
|
||||
through the new `Session.chat()`:
|
||||
- `diag_session_nemotron.py` — cross-process hamster recall: ✅
|
||||
- `mood_batch_poc.py` — cross-process mood trajectory: ✅
|
||||
- `mood_stream_poc.py` — 15 streaming turns + cross-process: ✅
|
||||
- `recall_poc.py` — per-project digest from 30-day git log: ✅ accurate,
|
||||
no hallucinated projects.
|
||||
|
||||
Falcon-Mamba was rerun through the new `Session.chat()` but still fails on
|
||||
the same tests — the trained model genuinely doesn't have the cross-turn
|
||||
recall capability that the hybrid Nemotron does. The 4 attention layers
|
||||
make the difference.
|
||||
|
||||
Reference in New Issue
Block a user