---
title: "Building Mentis: A Memory Layer for Coding Agents"
description: "How Mentis stores coding-agent attempts, retrieves related tasks, and marks outdated conclusions."
date: 2026-09-26T00:00:00.000Z
---

import MentisDiagram from '../../components/blog-visuals/MentisDiagram.astro';

I built Mentis because I wanted coding agents to reuse earlier investigations without assuming the results still apply. If an agent ran a test last week, that result is useful context, but it does not tell us whether the test passes now. Mentis stores the action, observation, and check so the next agent can find the attempt and verify it against the current code.

## Architecture and responsibilities

Mentis is an MCP tool server backed by Neo4j. It exposes `search`, `recall`, `record_attempt`, `mark_conclusion_outdated`, and `forget_attempt`. The external coding agent selects the tools, evaluates retrieved records, and runs checks. Mentis does not inspect the current checkout or modify code.

The graph has three levels: Repository → Task → Attempt. A task is identified by repository identity and task ID. An attempt stores an action, an observation, and optionally an inference and a check. If no check is provided, its status is `unverified`. A passed check applies only to the code tested during that attempt.

<MentisDiagram scene="system" />

### MCP transports

The local Node server communicates over stdio and checks Neo4j connectivity at startup. An optional Worker accepts POST requests at `/mcp` and creates and closes a server and database connection for each request. Both transports register the same strict tool schemas. Agents do not communicate directly through these transports; they share records in the graph.

<MentisDiagram scene="mcp" />

The documented Worker implementation has no caller authentication. Managed read transactions and response limits do not make agent-authored Cypher safe for untrusted callers. A deployment with access to a sensitive graph requires access controls and restricted database credentials.

## Search for related tasks

`search` accepts a natural-language problem description. OpenRouter embeds the query with `voyageai/voyage-4`. Neo4j's separately created `attempt_embedding` index returns up to 200 similar attempts. Mentis groups the results by `(repository, taskId)`, keeps at most five attempts per task, and uses `typesafe/jev-1.13` to score usefulness. If any Jev score fails, the whole search falls back to vector similarity. Embedding and vector-index failures return tool errors.

The result is a list of tasks, each with one matched attempt preview. Similarity and usefulness scores do not indicate correctness. Relevant attempts outside the vector result pool may be omitted. Returned attempts may include failed checks or outdated conclusions.

<MentisDiagram scene="search" />

### Read task history

`recall` runs agent-supplied Cypher to read attempts directly, without OpenRouter or a vector index. Queries should normally use both repository and task ID as parameters. The tool returns columns and rows as JSON text in the MCP result. It flags results truncated at 100 rows or 512,000 serialized bytes, and has a five-second read timeout. These limits do not bound the cost of every possible query.

<MentisDiagram scene="recall" />

Include failed attempts and correction fields in the Cypher projection when reviewing task history. Verify relevant observations against the current checkout before using them.

## Record an attempt

The agent calls `record_attempt` with the repository, task ID, code context, action, affected files, and observed result. An optional inference is stored separately from the observation. The agent should report a check as `passed` or `failed` only if it ran the check.

Mentis embeds the document text before starting the Neo4j write transaction. If embedding fails or does not return exactly 1,024 finite numbers, nothing is written. On success, one transaction merges the repository and task and adds a new attempt with a UUID. Subsequent calls add attempts rather than replacing earlier records.

<MentisDiagram scene="record" />

## Mark an outdated conclusion

`mark_conclusion_outdated` matches an attempt by repository and attempt ID. It stores a reason and timestamp and can also store an agent-reported latest commit. The original action, observation, check result, and embedding remain unchanged. Search can still return the attempt with its correction. Recall queries must explicitly request the correction properties.

Calling the tool again replaces the previous reason and timestamp; it does not preserve a correction history. New evidence should be stored as a new attempt. `forget_attempt` deletes an attempt from the live graph and is separate from marking a conclusion outdated.

<MentisDiagram scene="correction" />

## Data handling and verification

Query text and attempt content can be sent to OpenRouter for embedding or scoring. Do not include secrets. Git commit and dirty-state fields are supplied by the agent and are not verified by Mentis. Before using a retrieved attempt, inspect its history and check its observations against the current files.