Files
emilandClaude Opus 4.7 75c9ee4576 Add memba MVP: C++ core, Python SDK, CLI, examples, experiments
C++ core (libmemba.so):
- include/memba/state.h — C API (state_new/free/save/load/get_size)
- src/state.cpp — MEMB file format: magic, version, SHA-256 model_id,
  CRC-32, opaque llama_state_*_data() blob
- src/cli.cpp — minimal demo binary with greedy sampler
- CMakeLists.txt + build.sh with llama.cpp submodule, CUDA auto-detect

Python SDK (memba):
- core.py — file I/O via llama-cpp-python's exposed C functions,
  unwraps _LlamaContext to access raw context pointer (≥0.3.x)
- session.py — high-level Session with auto-save/load, ChatML wrapper
  for instruct models, raw mode for base models
- cli.py — typer-based: chat (REPL), run (one-shot), list, rm, info

Examples:
- 01_basic_save_load.py, 02_chat_session.py

Experiments (throwaway POCs documenting product-direction findings):
- recall_poc.py — git log → state → cross-process query
- mood_poc.py — batch sentiment trajectory, Mamba vs Transformer
- mood_stream_poc.py, mood_batch_poc.py — variants
- diag_saveload.py — minimal save/load isolation test
- README.md documents the headline finding: save/load is byte-identical,
  but Falcon-Mamba-7B-Instruct does not retain facts across conversation
  turns even in-process — limits viable products to single-prompt analysis
  and persona priming.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-16 12:48:37 +03:00

300 lines
8.2 KiB
Markdown

# memba
> Save, load and share model understanding — not weights, not chat history, but accumulated intelligence.
**memba** is a persistent memory layer for **SSM-based LLMs** (Falcon-Mamba, Zamba and other
recurrent architectures). It snapshots the model's internal hidden state to a compact binary
file so you can resume, branch, or share a "trained context" across processes, machines, or time.
```
First session Second session (different process / machine)
──────────────── ───────────────────────────────────────────
Feed 10 000 tokens of Load state file →
research papers → Ask follow-up question →
Save state file Model answers as if it just read those papers
```
**memba stores the SSM state, not the source text.** Privacy is structural: the raw documents
you processed are never written to disk by memba.
---
## Scope (MVP)
| In scope | Out of scope |
|----------|-------------|
| Falcon-Mamba, Zamba (SSM/Mamba architecture) | Transformer / KV-cache models |
| CPU ↔ GPU portable state files | Cloud sync, encryption |
| Python high-level API + CLI | Web scraping, RAG pipeline |
| C API + shared library | Dataset generation |
| Linux (primary), macOS (best-effort) | Windows |
---
## Installation
### Prerequisites
- CMake ≥ 3.14, a C++17 compiler (GCC ≥ 10 or Clang ≥ 12)
- Python ≥ 3.10
- *(Optional)* CUDA toolkit for GPU offload
### 1. Clone with submodule
```bash
git clone --recurse-submodules https://github.com/your-org/memba.git
cd memba
```
Or if you already cloned:
```bash
git submodule update --init --recursive
```
### 2. Build the C++ library and CLI
```bash
./build.sh # auto-detects CUDA / Metal
# or pass cmake flags directly:
./build.sh -DGGML_CUDA=ON
```
Artifacts:
- `build/libmemba.so` — shared library for C/C++ integration
- `build/memba-cli` — CLI demo binary
### 3. Install the Python package
```bash
pip install -e . # editable install, uses build/ for libmemba.so
# or standard install after building:
pip install .
```
The Python layer uses **llama-cpp-python** for inference and calls its embedded `libllama.so`
directly — no ABI conflict with your own `libmemba.so` build.
---
## Usage
### Python — high-level Session API
```python
from memba import Session
# CPU
s = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
print(s.chat("The transformer architecture was introduced in 2017 by Vaswani et al."))
s.save() # writes ~/.memba/states/research.memb
# Another process or the next day
s2 = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
print(s2.chat("Who were the authors?")) # model has context of prior statement
s2.save()
```
```python
# GPU offload
s = Session(
"falcon-mamba-7b-Q4_K_M.gguf",
session_id="gpu_session",
n_gpu_layers=-1, # -1 = all layers
n_ctx=8192,
)
print(s.chat("Explain quantum entanglement."))
print(f"State size: {s.state_size:,} bytes")
s.save()
```
### Python — low-level core API
```python
from llama_cpp import Llama
from memba import core
llama = Llama("falcon-mamba-7b-Q4_K_M.gguf", n_ctx=4096)
# Run some inference…
llama("The capital of France is Paris.", max_tokens=1)
# Checkpoint
core.save_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")
# … later / elsewhere …
core.load_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")
out = llama(" Its population is", max_tokens=32, echo=False)
print(out["choices"][0]["text"])
```
### Python CLI
```bash
# Interactive REPL (auto-saves on exit)
memba chat --model falcon-mamba-7b-Q4_K_M.gguf --session my_research
# One-shot with explicit state management
memba run --model falcon-mamba-7b-Q4_K_M.gguf \
--prompt "Capital of France is" \
--save-state /tmp/paris.memb
memba run --model falcon-mamba-7b-Q4_K_M.gguf \
--load-state /tmp/paris.memb \
--prompt " Its population is"
# Session management
memba list
memba info my_research
memba rm old_session
```
### C++ CLI
```bash
# CPU — save state after generation
./build/memba-cli \
--model falcon-mamba-7b-Q4_K_M.gguf \
--prompt "Capital of France is" \
--save-state paris.bin
# CPU — load state and continue
./build/memba-cli \
--model falcon-mamba-7b-Q4_K_M.gguf \
--load-state paris.bin \
--prompt " Its population is"
# GPU
./build/memba-cli \
--model falcon-mamba-7b-Q4_K_M.gguf \
--n-gpu-layers 35 \
--prompt "Hello" \
--save-state gpu_session.bin
```
### C API
```c
#include <memba/state.h>
#include <llama.h>
llama_model* model = llama_model_load_from_file("model.gguf", llama_model_default_params());
llama_context* ctx = llama_new_context_with_model(model, llama_context_default_params());
memba_state_t* state = memba_state_new(ctx, "model.gguf");
// … run inference …
int rc = memba_state_save(state, "checkpoint.memb");
if (rc != MEMBA_OK) fprintf(stderr, "%s\n", memba_error_string(rc));
// … later …
rc = memba_state_load(state, "checkpoint.memb");
memba_state_free(state);
llama_free(ctx);
llama_model_free(model);
```
---
## State file format
```
Offset Size Field
──────────────────────────────────────────────────────────────────
0 4 Magic: "MEMB"
4 4 Version: uint32 (1)
8 64 model_id: SHA-256 hex of first 1 KiB of GGUF (ASCII)
72 4 n_ctx: uint32
76 4 llama_ver: uint32 (reserved, 0)
80 8 data_size: uint64
88 N Opaque SSM state blob (llama_state_get_data output)
88+N 4 CRC-32 of the blob (IEEE 802.3 polynomial)
```
All integers are **little-endian**.
The blob is completely opaque — memba never parses its internals.
The `model_id` field prevents accidentally loading a state into the wrong model.
---
## Limitations
- **SSM models only.** Transformer KV-cache is orders of magnitude larger and architecturally
incompatible with this approach.
- **Same llama.cpp version required.** The opaque blob format can change between llama.cpp
builds. Pin your llama.cpp submodule commit when sharing state files across machines.
- **Same model file required.** The `model_id` check compares SHA-256 of the first 1 KiB of
the GGUF. Quantisation variants of the same base model will have different IDs.
- **No encryption.** The state file is unencrypted. Treat it with the same care as the model
weights.
- **No Windows support** in this MVP (path handling and shared-library loading not tested).
---
## Development
```bash
# Install dev dependencies
pip install -e ".[dev]"
# Lint
ruff check python/
# Type-check
mypy python/memba/
# Tests (requires a GGUF model — set MEMBA_TEST_MODEL env var)
pytest tests/ -v
```
---
## Project structure
```
memba/
├── llama.cpp/ git submodule (ggerganov/llama.cpp, MIT)
├── include/memba/
│ └── state.h C API (public header)
├── src/
│ ├── state.cpp C++ implementation of save/load
│ └── cli.cpp C++ CLI demo
├── python/memba/
│ ├── __init__.py
│ ├── core.py Low-level state I/O (ctypes → llama-cpp-python)
│ ├── session.py High-level Session class
│ └── cli.py typer CLI (memba chat / run / list / rm / info)
├── examples/
│ ├── 01_basic_save_load.py
│ └── 02_chat_session.py
├── CMakeLists.txt
├── pyproject.toml
├── build.sh
└── README.md
```
---
## Roadmap
- [ ] `memba fork <session> <new-name>` — branch a state for parallel exploration
- [ ] State diff / merge (experimental)
- [ ] Encryption at rest (AES-256-GCM)
- [ ] Cloud sync backend (S3-compatible)
- [ ] Dataset generation from accumulated states
---
## License
MIT — see [LICENSE](LICENSE).
---
## Credits
Built on **[llama.cpp](https://github.com/ggerganov/llama.cpp)** by Georgi Gerganov and
contributors (MIT). The core state serialisation primitives (`llama_state_get_data` /
`llama_state_set_data`) are part of llama.cpp's public API.