C++ core (libmemba.so): - include/memba/state.h — C API (state_new/free/save/load/get_size) - src/state.cpp — MEMB file format: magic, version, SHA-256 model_id, CRC-32, opaque llama_state_*_data() blob - src/cli.cpp — minimal demo binary with greedy sampler - CMakeLists.txt + build.sh with llama.cpp submodule, CUDA auto-detect Python SDK (memba): - core.py — file I/O via llama-cpp-python's exposed C functions, unwraps _LlamaContext to access raw context pointer (≥0.3.x) - session.py — high-level Session with auto-save/load, ChatML wrapper for instruct models, raw mode for base models - cli.py — typer-based: chat (REPL), run (one-shot), list, rm, info Examples: - 01_basic_save_load.py, 02_chat_session.py Experiments (throwaway POCs documenting product-direction findings): - recall_poc.py — git log → state → cross-process query - mood_poc.py — batch sentiment trajectory, Mamba vs Transformer - mood_stream_poc.py, mood_batch_poc.py — variants - diag_saveload.py — minimal save/load isolation test - README.md documents the headline finding: save/load is byte-identical, but Falcon-Mamba-7B-Instruct does not retain facts across conversation turns even in-process — limits viable products to single-prompt analysis and persona priming. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
300 lines
8.2 KiB
Markdown
300 lines
8.2 KiB
Markdown
# memba
|
|
|
|
> Save, load and share model understanding — not weights, not chat history, but accumulated intelligence.
|
|
|
|
**memba** is a persistent memory layer for **SSM-based LLMs** (Falcon-Mamba, Zamba and other
|
|
recurrent architectures). It snapshots the model's internal hidden state to a compact binary
|
|
file so you can resume, branch, or share a "trained context" across processes, machines, or time.
|
|
|
|
```
|
|
First session Second session (different process / machine)
|
|
──────────────── ───────────────────────────────────────────
|
|
Feed 10 000 tokens of Load state file →
|
|
research papers → Ask follow-up question →
|
|
Save state file Model answers as if it just read those papers
|
|
```
|
|
|
|
**memba stores the SSM state, not the source text.** Privacy is structural: the raw documents
|
|
you processed are never written to disk by memba.
|
|
|
|
---
|
|
|
|
## Scope (MVP)
|
|
|
|
| In scope | Out of scope |
|
|
|----------|-------------|
|
|
| Falcon-Mamba, Zamba (SSM/Mamba architecture) | Transformer / KV-cache models |
|
|
| CPU ↔ GPU portable state files | Cloud sync, encryption |
|
|
| Python high-level API + CLI | Web scraping, RAG pipeline |
|
|
| C API + shared library | Dataset generation |
|
|
| Linux (primary), macOS (best-effort) | Windows |
|
|
|
|
---
|
|
|
|
## Installation
|
|
|
|
### Prerequisites
|
|
|
|
- CMake ≥ 3.14, a C++17 compiler (GCC ≥ 10 or Clang ≥ 12)
|
|
- Python ≥ 3.10
|
|
- *(Optional)* CUDA toolkit for GPU offload
|
|
|
|
### 1. Clone with submodule
|
|
|
|
```bash
|
|
git clone --recurse-submodules https://github.com/your-org/memba.git
|
|
cd memba
|
|
```
|
|
|
|
Or if you already cloned:
|
|
|
|
```bash
|
|
git submodule update --init --recursive
|
|
```
|
|
|
|
### 2. Build the C++ library and CLI
|
|
|
|
```bash
|
|
./build.sh # auto-detects CUDA / Metal
|
|
# or pass cmake flags directly:
|
|
./build.sh -DGGML_CUDA=ON
|
|
```
|
|
|
|
Artifacts:
|
|
- `build/libmemba.so` — shared library for C/C++ integration
|
|
- `build/memba-cli` — CLI demo binary
|
|
|
|
### 3. Install the Python package
|
|
|
|
```bash
|
|
pip install -e . # editable install, uses build/ for libmemba.so
|
|
# or standard install after building:
|
|
pip install .
|
|
```
|
|
|
|
The Python layer uses **llama-cpp-python** for inference and calls its embedded `libllama.so`
|
|
directly — no ABI conflict with your own `libmemba.so` build.
|
|
|
|
---
|
|
|
|
## Usage
|
|
|
|
### Python — high-level Session API
|
|
|
|
```python
|
|
from memba import Session
|
|
|
|
# CPU
|
|
s = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
|
|
print(s.chat("The transformer architecture was introduced in 2017 by Vaswani et al."))
|
|
s.save() # writes ~/.memba/states/research.memb
|
|
|
|
# Another process or the next day
|
|
s2 = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
|
|
print(s2.chat("Who were the authors?")) # model has context of prior statement
|
|
s2.save()
|
|
```
|
|
|
|
```python
|
|
# GPU offload
|
|
s = Session(
|
|
"falcon-mamba-7b-Q4_K_M.gguf",
|
|
session_id="gpu_session",
|
|
n_gpu_layers=-1, # -1 = all layers
|
|
n_ctx=8192,
|
|
)
|
|
print(s.chat("Explain quantum entanglement."))
|
|
print(f"State size: {s.state_size:,} bytes")
|
|
s.save()
|
|
```
|
|
|
|
### Python — low-level core API
|
|
|
|
```python
|
|
from llama_cpp import Llama
|
|
from memba import core
|
|
|
|
llama = Llama("falcon-mamba-7b-Q4_K_M.gguf", n_ctx=4096)
|
|
|
|
# Run some inference…
|
|
llama("The capital of France is Paris.", max_tokens=1)
|
|
|
|
# Checkpoint
|
|
core.save_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")
|
|
|
|
# … later / elsewhere …
|
|
core.load_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")
|
|
out = llama(" Its population is", max_tokens=32, echo=False)
|
|
print(out["choices"][0]["text"])
|
|
```
|
|
|
|
### Python CLI
|
|
|
|
```bash
|
|
# Interactive REPL (auto-saves on exit)
|
|
memba chat --model falcon-mamba-7b-Q4_K_M.gguf --session my_research
|
|
|
|
# One-shot with explicit state management
|
|
memba run --model falcon-mamba-7b-Q4_K_M.gguf \
|
|
--prompt "Capital of France is" \
|
|
--save-state /tmp/paris.memb
|
|
|
|
memba run --model falcon-mamba-7b-Q4_K_M.gguf \
|
|
--load-state /tmp/paris.memb \
|
|
--prompt " Its population is"
|
|
|
|
# Session management
|
|
memba list
|
|
memba info my_research
|
|
memba rm old_session
|
|
```
|
|
|
|
### C++ CLI
|
|
|
|
```bash
|
|
# CPU — save state after generation
|
|
./build/memba-cli \
|
|
--model falcon-mamba-7b-Q4_K_M.gguf \
|
|
--prompt "Capital of France is" \
|
|
--save-state paris.bin
|
|
|
|
# CPU — load state and continue
|
|
./build/memba-cli \
|
|
--model falcon-mamba-7b-Q4_K_M.gguf \
|
|
--load-state paris.bin \
|
|
--prompt " Its population is"
|
|
|
|
# GPU
|
|
./build/memba-cli \
|
|
--model falcon-mamba-7b-Q4_K_M.gguf \
|
|
--n-gpu-layers 35 \
|
|
--prompt "Hello" \
|
|
--save-state gpu_session.bin
|
|
```
|
|
|
|
### C API
|
|
|
|
```c
|
|
#include <memba/state.h>
|
|
#include <llama.h>
|
|
|
|
llama_model* model = llama_model_load_from_file("model.gguf", llama_model_default_params());
|
|
llama_context* ctx = llama_new_context_with_model(model, llama_context_default_params());
|
|
memba_state_t* state = memba_state_new(ctx, "model.gguf");
|
|
|
|
// … run inference …
|
|
|
|
int rc = memba_state_save(state, "checkpoint.memb");
|
|
if (rc != MEMBA_OK) fprintf(stderr, "%s\n", memba_error_string(rc));
|
|
|
|
// … later …
|
|
rc = memba_state_load(state, "checkpoint.memb");
|
|
|
|
memba_state_free(state);
|
|
llama_free(ctx);
|
|
llama_model_free(model);
|
|
```
|
|
|
|
---
|
|
|
|
## State file format
|
|
|
|
```
|
|
Offset Size Field
|
|
──────────────────────────────────────────────────────────────────
|
|
0 4 Magic: "MEMB"
|
|
4 4 Version: uint32 (1)
|
|
8 64 model_id: SHA-256 hex of first 1 KiB of GGUF (ASCII)
|
|
72 4 n_ctx: uint32
|
|
76 4 llama_ver: uint32 (reserved, 0)
|
|
80 8 data_size: uint64
|
|
88 N Opaque SSM state blob (llama_state_get_data output)
|
|
88+N 4 CRC-32 of the blob (IEEE 802.3 polynomial)
|
|
```
|
|
|
|
All integers are **little-endian**.
|
|
The blob is completely opaque — memba never parses its internals.
|
|
The `model_id` field prevents accidentally loading a state into the wrong model.
|
|
|
|
---
|
|
|
|
## Limitations
|
|
|
|
- **SSM models only.** Transformer KV-cache is orders of magnitude larger and architecturally
|
|
incompatible with this approach.
|
|
- **Same llama.cpp version required.** The opaque blob format can change between llama.cpp
|
|
builds. Pin your llama.cpp submodule commit when sharing state files across machines.
|
|
- **Same model file required.** The `model_id` check compares SHA-256 of the first 1 KiB of
|
|
the GGUF. Quantisation variants of the same base model will have different IDs.
|
|
- **No encryption.** The state file is unencrypted. Treat it with the same care as the model
|
|
weights.
|
|
- **No Windows support** in this MVP (path handling and shared-library loading not tested).
|
|
|
|
---
|
|
|
|
## Development
|
|
|
|
```bash
|
|
# Install dev dependencies
|
|
pip install -e ".[dev]"
|
|
|
|
# Lint
|
|
ruff check python/
|
|
|
|
# Type-check
|
|
mypy python/memba/
|
|
|
|
# Tests (requires a GGUF model — set MEMBA_TEST_MODEL env var)
|
|
pytest tests/ -v
|
|
```
|
|
|
|
---
|
|
|
|
## Project structure
|
|
|
|
```
|
|
memba/
|
|
├── llama.cpp/ git submodule (ggerganov/llama.cpp, MIT)
|
|
├── include/memba/
|
|
│ └── state.h C API (public header)
|
|
├── src/
|
|
│ ├── state.cpp C++ implementation of save/load
|
|
│ └── cli.cpp C++ CLI demo
|
|
├── python/memba/
|
|
│ ├── __init__.py
|
|
│ ├── core.py Low-level state I/O (ctypes → llama-cpp-python)
|
|
│ ├── session.py High-level Session class
|
|
│ └── cli.py typer CLI (memba chat / run / list / rm / info)
|
|
├── examples/
|
|
│ ├── 01_basic_save_load.py
|
|
│ └── 02_chat_session.py
|
|
├── CMakeLists.txt
|
|
├── pyproject.toml
|
|
├── build.sh
|
|
└── README.md
|
|
```
|
|
|
|
---
|
|
|
|
## Roadmap
|
|
|
|
- [ ] `memba fork <session> <new-name>` — branch a state for parallel exploration
|
|
- [ ] State diff / merge (experimental)
|
|
- [ ] Encryption at rest (AES-256-GCM)
|
|
- [ ] Cloud sync backend (S3-compatible)
|
|
- [ ] Dataset generation from accumulated states
|
|
|
|
---
|
|
|
|
## License
|
|
|
|
MIT — see [LICENSE](LICENSE).
|
|
|
|
---
|
|
|
|
## Credits
|
|
|
|
Built on **[llama.cpp](https://github.com/ggerganov/llama.cpp)** by Georgi Gerganov and
|
|
contributors (MIT). The core state serialisation primitives (`llama_state_get_data` /
|
|
`llama_state_set_data`) are part of llama.cpp's public API.
|