Two bugs were blocking memba's main promise (load .memb in a fresh process → model continues with full recalled context): 1. llama_state_set_data() restores the C-level KV-cache + SSM hidden state, but llama-cpp-python's Python wrapper still reports n_tokens=0. The next eval() then decodes new tokens at offset 0 and overwrites the loaded state. Fix: extend MEMB format with an optional 12-byte trailer appended after the CRC32. It carries the wrapper's n_tokens. The C library reads up to CRC and ignores anything past it, so files stay backward-compatible with libmemba; only the Python loader uses it. 2. Llama.__call__ / create_chat_completion / generate all re-tokenise the prompt on every call and clear the KV-cache when the new tokens don't prefix-match input_ids. That destroys any state we just loaded. Fix: rewrite Session.chat() to use raw tokenize → eval → sample. eval() appends tokens to the live state without resetting, and we handle stop-token detection ourselves. Verified end-to-end on Nemotron-3-Nano-4B (hybrid 21x Mamba-2 + 4x attention) — see experiments/README.md for the full findings log. diag_session_nemotron.py, mood_batch_poc.py, mood_stream_poc.py and recall_poc.py now all pass their cross-process tests; Falcon-Mamba still fails because the trained model itself can't do cross-turn recall — that was the original misdiagnosis. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
70 lines
3.0 KiB
Python
70 lines
3.0 KiB
Python
"""
|
|
mood_batch_poc.py — batch ingest, cross-process query.
|
|
|
|
This is the clean test: one chat() call with all 15 messages as a block,
|
|
save state, EXIT, then in a fresh process load state and ask sentiment
|
|
questions. Isolates the cross-process save/load from streaming-noise.
|
|
"""
|
|
from __future__ import annotations
|
|
import sys, argparse
|
|
from pathlib import Path
|
|
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "python"))
|
|
from memba import Session
|
|
|
|
MAMBA = "/home/emil/Desktop/Coding/AI/Memba/NVIDIA-Nemotron3-Nano-4B-Q4_K_M.gguf"
|
|
STATE_DIR = "/tmp/mood_batch_test"
|
|
SESSION = "mood_batch"
|
|
|
|
CHAT_LOG = [
|
|
"Morning team! Coffee in hand, ready to tackle the auth refactor today.",
|
|
"Just pushed PR #234 fixing the token validation bug. Should be a quick merge.",
|
|
"Code review comments came in fast, all good catches. Iterating now.",
|
|
"Basic flow working locally, tests passing. Feeling good about this.",
|
|
"Heading to lunch, hopefully wrap this up by EOD.",
|
|
"Back. CI is failing on something unrelated, looking into it.",
|
|
"OK the 'unrelated' thing is actually related. Auth tests use a stale fixture.",
|
|
"Why does the fixture rebuild take 12 minutes. Every. Single. Time.",
|
|
"Cancelled the run twice now. Going to bypass and run tests locally.",
|
|
"Local passes, CI fails. Classic.",
|
|
"Two hours gone on this fixture issue. Not even what I was supposed to be doing.",
|
|
"Now there's a merge conflict with main because someone restructured migrations.",
|
|
"Whoever shipped those migrations on a Friday afternoon, I will find you.",
|
|
"Closing the laptop. Will fight this tomorrow.",
|
|
"Actually no. One more try before I sleep.",
|
|
]
|
|
|
|
INGEST_PROMPT = (
|
|
"You are observing one person's chat messages from a workday. "
|
|
"Here they are in order. Read them and remember the overall trajectory. "
|
|
"Reply with just 'noted'.\n\n"
|
|
+ "\n".join(f"[msg {i+1:>2}] {m}" for i, m in enumerate(CHAT_LOG))
|
|
)
|
|
|
|
|
|
def build():
|
|
p = Path(STATE_DIR) / f"{SESSION}.memb"
|
|
if p.exists(): p.unlink()
|
|
s = Session(model_path=MAMBA, session_id=SESSION, state_dir=STATE_DIR,
|
|
n_gpu_layers=-1, n_ctx=4096, chat_format="chatml")
|
|
print(f"[build] ack: {s.chat(INGEST_PROMPT, max_tokens=8)!r}")
|
|
print(f"[build] state: {s.state_size:,} B")
|
|
s.save()
|
|
|
|
|
|
def query():
|
|
s = Session(model_path=MAMBA, session_id=SESSION, state_dir=STATE_DIR,
|
|
n_gpu_layers=-1, n_ctx=4096, chat_format="chatml")
|
|
print(f"[query] loaded {s.state_size:,} B\n")
|
|
for q in [
|
|
"What is this person's current emotional state? One sentence.",
|
|
"Did their mood change over the messages? One sentence describing the trajectory.",
|
|
"Around which message number did the mood shift from positive to negative? Just the number.",
|
|
]:
|
|
print(f"[Q] {q}")
|
|
print(f"[A] {s.chat(q, max_tokens=120)}\n")
|
|
|
|
|
|
if __name__ == "__main__":
|
|
cmd = sys.argv[1] if len(sys.argv) > 1 else "build"
|
|
{"build": build, "query": query}[cmd]()
|