SOKKAN 3.0 “One memory”
SOKKAN’s memory.
What it does, what it is worth once measured, which machine it runs on, and the licence of the model that powers it. This page also says what we do not know yet.
The principle
Your agents write notes: decisions, conventions, procedures, pitfalls. SOKKAN keeps them in one memory, which the CortHeXis tab shows as a graph and keeps reviewing.
- When a session starts, the project’s relevant notes are already in its context.
- With every message, SOKKAN looks up what the question touches and adds it, with the age of each note and where that date comes from (“updated 23 days ago, reconstructed date”).
- For every sub-agent, the recall for its topic is added to its instructions.
- Every recall is logged: which session or sub-agent received which note, in which version.
- Recalled notes are presented to the agent as data, not instructions.
The CortHeXis review
A health score from 0 to 100 and its history; findings grouped (chain, structure, drift, security) with their fix; one-click actions, all subject to approval: relink, merge two notes with a diff, rename to the convention, close a dormant work item. Cases that need judgement open a curation session preloaded with the findings. Alerts via Telegram or webhook.
The benchmark: what the memory is worth, measured
We asked 300 questions (195 in French, 75 in English, 30 in German, 52 of which are answered only in the body of a note) of our own working memory: 415 notes, 2,498 passages, on 2–3 October 2026. For each question, the expected note is known.
| Configuration | Right note first | MRR |
|---|---|---|
| SOKKAN 2.x | 42% | 0.55 |
| SOKKAN 3.0, Light and Standard profiles (no GPU) | 73% | 0.82 |
| SOKKAN 3.0, GPU or ANCHOR profile (second ranking on the top 10) | 82% | 0.88 |
| SOKKAN 3.0 with the MIT-licensed fallback model | not measured | 0.75 to 0.77 |
How to read it. “Right note first”: the expected note is the first result. MRR (mean reciprocal rank): 1 if the right note always comes first, 0.5 if it comes second on average. The no-GPU profiles were measured in containers capped at 4 cores / 4 GB and 8 cores / 16 GB.
What we do not publish. The corpus itself: it is our real working memory and holds our business. That is why the benchmark ships with SOKKAN: it measures the memory on your notes, with questions harvested from your own sessions, and it blocks any profile switch that would do worse.
Still to be measured. Our figures come from a 415-note memory. At 250,000 passages, vector search itself stays under 30 ms in our measurements, but ranking quality depends on your content: that is exactly what the built-in benchmark will tell you.
Profiles, picked by Magnitude
| Profile | Machine | What changes |
|---|---|---|
| Light | 4 cores, 4 GB, no GPU | Hybrid search (meaning and keywords), no second ranking. Everything stays on the machine. |
| Standard | 8 cores, 16 GB, no GPU | Like Light, plus a finer second ranking in the background (deferred recall, digest): it costs several seconds per search on CPU, hence deferred. |
| GPU · ANCHOR | A GPU, or the ANCHOR appliance | Second ranking on every search; importing a large memory takes hours rather than nights (250,000 passages: about 1 h 25 on GPU against 14 to 20 h on CPU, from measured throughput). |
Switching profile builds a new index generation in the background while the old one keeps answering. The switch only happens if the benchmark does not regress, and the old generation is kept for 7 days for rollback. If the model changes, the index is rebuilt in full: two models are never mixed in the same index.
The memory model’s licence
In one sentence: SOKKAN stays open source (Apache-2.0); the model behind the memory, Google’s EmbeddingGemma, is not. So SOKKAN does not ship it: it downloads it on first launch, once you have read and accepted Google’s terms. If you decline, memory works with an MIT-licensed model.
Why this model
It is the best one we measured on our benchmark, and it is light: it runs without a GPU on a small server, in French, English and German. It finds the right note far more often than SOKKAN 2.x’s model (see the benchmark).
What Google’s terms say
- Commercial use is allowed, including in a paid product, on premises or in the cloud.
- It is not an open-source licence: Google imposes a prohibited use policy (illegal activities, privacy violations, disinformation, etc.) and may change it.
- Google reserves the right to restrict use it considers contrary to these terms.
- Google claims no rights over the outputs: the computed vectors are yours.
Official texts: Gemma Terms of Use and the prohibited use policy.
What SOKKAN does
- The SOKKAN repository does not contain the model weights: the code stays entirely Apache-2.0.
- On first launch, the Magnitude screen shows the model, its publisher, the download size (about 330 MB), a summary of the terms and the official links, then lets you choose: “I accept and download” or “Use the free model (MIT)”. The acceptance (date, terms version, who) is written to the audit log.
- The file is checked against its fingerprint before use. The model runs on your side with no call to Google: your notes never leave the machine.
- Offline or if you decline, SOKKAN uses the MIT model with no loss of features (and slightly weaker recall). You can change your mind later: the index is rebuilt in the background.
- On SOKKAN Cloud and on the ANCHOR appliance, we run the model; our terms carry Google’s use policy and the required notice: “Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms”.
Coming from SOKKAN 2.x
- Your notes stay untouched. Before any normalisation, the corpus is archived, and the proposed changes (renames, merges, repaired headers) are shown to you before they are applied.
- The old index is not converted: memory is reindexed in the background, and search keeps answering from the old index in the meantime.
- The memory tools offered to agents keep the same names and arguments; their results gain the age and source of dates.
- SOKKAN Cloud instances are migrated automatically.