emilandClaude Opus 4.7 75c9ee4576 Add memba MVP: C++ core, Python SDK, CLI, examples, experiments
C++ core (libmemba.so):
- include/memba/state.h — C API (state_new/free/save/load/get_size)
- src/state.cpp — MEMB file format: magic, version, SHA-256 model_id,
  CRC-32, opaque llama_state_*_data() blob
- src/cli.cpp — minimal demo binary with greedy sampler
- CMakeLists.txt + build.sh with llama.cpp submodule, CUDA auto-detect

Python SDK (memba):
- core.py — file I/O via llama-cpp-python's exposed C functions,
  unwraps _LlamaContext to access raw context pointer (≥0.3.x)
- session.py — high-level Session with auto-save/load, ChatML wrapper
  for instruct models, raw mode for base models
- cli.py — typer-based: chat (REPL), run (one-shot), list, rm, info

Examples:
- 01_basic_save_load.py, 02_chat_session.py

Experiments (throwaway POCs documenting product-direction findings):
- recall_poc.py — git log → state → cross-process query
- mood_poc.py — batch sentiment trajectory, Mamba vs Transformer
- mood_stream_poc.py, mood_batch_poc.py — variants
- diag_saveload.py — minimal save/load isolation test
- README.md documents the headline finding: save/load is byte-identical,
  but Falcon-Mamba-7B-Instruct does not retain facts across conversation
  turns even in-process — limits viable products to single-prompt analysis
  and persona priming.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-16 12:48:37 +03:00
2026-05-16 07:01:57 +03:00

memba

Save, load and share model understanding — not weights, not chat history, but accumulated intelligence.

memba is a persistent memory layer for SSM-based LLMs (Falcon-Mamba, Zamba and other recurrent architectures). It snapshots the model's internal hidden state to a compact binary file so you can resume, branch, or share a "trained context" across processes, machines, or time.

First session               Second session (different process / machine)
────────────────            ───────────────────────────────────────────
Feed 10 000 tokens of       Load state file →
research papers        →    Ask follow-up question →
Save state file             Model answers as if it just read those papers

memba stores the SSM state, not the source text. Privacy is structural: the raw documents you processed are never written to disk by memba.


Scope (MVP)

In scope Out of scope
Falcon-Mamba, Zamba (SSM/Mamba architecture) Transformer / KV-cache models
CPU ↔ GPU portable state files Cloud sync, encryption
Python high-level API + CLI Web scraping, RAG pipeline
C API + shared library Dataset generation
Linux (primary), macOS (best-effort) Windows

Installation

Prerequisites

  • CMake ≥ 3.14, a C++17 compiler (GCC ≥ 10 or Clang ≥ 12)
  • Python ≥ 3.10
  • (Optional) CUDA toolkit for GPU offload

1. Clone with submodule

git clone --recurse-submodules https://github.com/your-org/memba.git
cd memba

Or if you already cloned:

git submodule update --init --recursive

2. Build the C++ library and CLI

./build.sh                          # auto-detects CUDA / Metal
# or pass cmake flags directly:
./build.sh -DGGML_CUDA=ON

Artifacts:

  • build/libmemba.so — shared library for C/C++ integration
  • build/memba-cli — CLI demo binary

3. Install the Python package

pip install -e .                    # editable install, uses build/ for libmemba.so
# or standard install after building:
pip install .

The Python layer uses llama-cpp-python for inference and calls its embedded libllama.so directly — no ABI conflict with your own libmemba.so build.


Usage

Python — high-level Session API

from memba import Session

# CPU
s = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
print(s.chat("The transformer architecture was introduced in 2017 by Vaswani et al."))
s.save()                            # writes ~/.memba/states/research.memb

# Another process or the next day
s2 = Session("falcon-mamba-7b-Q4_K_M.gguf", session_id="research")
print(s2.chat("Who were the authors?"))   # model has context of prior statement
s2.save()
# GPU offload
s = Session(
    "falcon-mamba-7b-Q4_K_M.gguf",
    session_id="gpu_session",
    n_gpu_layers=-1,                # -1 = all layers
    n_ctx=8192,
)
print(s.chat("Explain quantum entanglement."))
print(f"State size: {s.state_size:,} bytes")
s.save()

Python — low-level core API

from llama_cpp import Llama
from memba import core

llama = Llama("falcon-mamba-7b-Q4_K_M.gguf", n_ctx=4096)

# Run some inference…
llama("The capital of France is Paris.", max_tokens=1)

# Checkpoint
core.save_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")

# … later / elsewhere …
core.load_state(llama, "falcon-mamba-7b-Q4_K_M.gguf", "/tmp/paris.memb")
out = llama(" Its population is", max_tokens=32, echo=False)
print(out["choices"][0]["text"])

Python CLI

# Interactive REPL (auto-saves on exit)
memba chat --model falcon-mamba-7b-Q4_K_M.gguf --session my_research

# One-shot with explicit state management
memba run  --model falcon-mamba-7b-Q4_K_M.gguf \
           --prompt "Capital of France is" \
           --save-state /tmp/paris.memb

memba run  --model falcon-mamba-7b-Q4_K_M.gguf \
           --load-state /tmp/paris.memb \
           --prompt " Its population is"

# Session management
memba list
memba info my_research
memba rm   old_session

C++ CLI

# CPU — save state after generation
./build/memba-cli \
    --model falcon-mamba-7b-Q4_K_M.gguf \
    --prompt "Capital of France is" \
    --save-state paris.bin

# CPU — load state and continue
./build/memba-cli \
    --model falcon-mamba-7b-Q4_K_M.gguf \
    --load-state paris.bin \
    --prompt " Its population is"

# GPU
./build/memba-cli \
    --model falcon-mamba-7b-Q4_K_M.gguf \
    --n-gpu-layers 35 \
    --prompt "Hello" \
    --save-state gpu_session.bin

C API

#include <memba/state.h>
#include <llama.h>

llama_model*   model = llama_model_load_from_file("model.gguf", llama_model_default_params());
llama_context* ctx   = llama_new_context_with_model(model, llama_context_default_params());
memba_state_t* state = memba_state_new(ctx, "model.gguf");

// … run inference …

int rc = memba_state_save(state, "checkpoint.memb");
if (rc != MEMBA_OK) fprintf(stderr, "%s\n", memba_error_string(rc));

// … later …
rc = memba_state_load(state, "checkpoint.memb");

memba_state_free(state);
llama_free(ctx);
llama_model_free(model);

State file format

Offset   Size   Field
──────────────────────────────────────────────────────────────────
0        4      Magic: "MEMB"
4        4      Version: uint32 (1)
8        64     model_id: SHA-256 hex of first 1 KiB of GGUF (ASCII)
72       4      n_ctx: uint32
76       4      llama_ver: uint32 (reserved, 0)
80       8      data_size: uint64
88       N      Opaque SSM state blob (llama_state_get_data output)
88+N     4      CRC-32 of the blob (IEEE 802.3 polynomial)

All integers are little-endian. The blob is completely opaque — memba never parses its internals. The model_id field prevents accidentally loading a state into the wrong model.


Limitations

  • SSM models only. Transformer KV-cache is orders of magnitude larger and architecturally incompatible with this approach.
  • Same llama.cpp version required. The opaque blob format can change between llama.cpp builds. Pin your llama.cpp submodule commit when sharing state files across machines.
  • Same model file required. The model_id check compares SHA-256 of the first 1 KiB of the GGUF. Quantisation variants of the same base model will have different IDs.
  • No encryption. The state file is unencrypted. Treat it with the same care as the model weights.
  • No Windows support in this MVP (path handling and shared-library loading not tested).

Development

# Install dev dependencies
pip install -e ".[dev]"

# Lint
ruff check python/

# Type-check
mypy python/memba/

# Tests (requires a GGUF model — set MEMBA_TEST_MODEL env var)
pytest tests/ -v

Project structure

memba/
├── llama.cpp/          git submodule (ggerganov/llama.cpp, MIT)
├── include/memba/
│   └── state.h         C API (public header)
├── src/
│   ├── state.cpp       C++ implementation of save/load
│   └── cli.cpp         C++ CLI demo
├── python/memba/
│   ├── __init__.py
│   ├── core.py         Low-level state I/O (ctypes → llama-cpp-python)
│   ├── session.py      High-level Session class
│   └── cli.py          typer CLI (memba chat / run / list / rm / info)
├── examples/
│   ├── 01_basic_save_load.py
│   └── 02_chat_session.py
├── CMakeLists.txt
├── pyproject.toml
├── build.sh
└── README.md

Roadmap

  • memba fork <session> <new-name> — branch a state for parallel exploration
  • State diff / merge (experimental)
  • Encryption at rest (AES-256-GCM)
  • Cloud sync backend (S3-compatible)
  • Dataset generation from accumulated states

License

MIT — see LICENSE.


Credits

Built on llama.cpp by Georgi Gerganov and contributors (MIT). The core state serialisation primitives (llama_state_get_data / llama_state_set_data) are part of llama.cpp's public API.

S
Description
No description provided
Readme
88 KiB
Languages
Python 67%
C++ 24.7%
C 3.8%
Shell 2.3%
CMake 2.2%