I built Mentis because I wanted coding agents to reuse earlier investigations without assuming the results still apply. If an agent ran a test last week, that result is useful context, but it does not tell us whether the test passes now. Mentis stores the action, observation, and check so the next agent can find the attempt and verify it against the current code.

Architecture and responsibilities

Mentis is an MCP tool server backed by Neo4j. It exposes search, recall, record_attempt, mark_conclusion_outdated, and forget_attempt. The external coding agent selects the tools, evaluates retrieved records, and runs checks. Mentis does not inspect the current checkout or modify code.

The graph has three levels: Repository → Task → Attempt. A task is identified by repository identity and task ID. An attempt stores an action, an observation, and optionally an inference and a check. If no check is provided, its status is unverified. A passed check applies only to the code tested during that attempt.

Where does the agent end and the memory layer begin?
Step 1 of 6 The agent investigates its checkout; memory does not inspect code.
Read all steps
  1. The agent investigates its checkout; memory does not inspect code.
  2. The agent chooses local stdio for this call (the Worker is an alternative).
  3. Both transports expose the same five validated tools.
  4. The tool delegates the requested operation to MemoryGraph.
  5. Search or recording sends text to OpenRouter; recall and correction do not.
  6. After embedding, search and recording use Neo4j; recall and correction reach it without OpenRouter.

MCP transports

The local Node server communicates over stdio and checks Neo4j connectivity at startup. An optional Worker accepts POST requests at /mcp and creates and closes a server and database connection for each request. Both transports register the same strict tool schemas. Agents do not communicate directly through these transports; they share records in the graph.

What messages cross the MCP boundary?
Step 1 of 7 Choose one transport. Local stdio verifies connectivity at startup; the Worker creates connections per POST.

Scroll sideways to follow all four actors.

Read all steps
  1. Choose one transport. Local stdio verifies connectivity at startup; the Worker creates connections per POST.
  2. The external agent sends tools/call with a tool name and JSON input.
  3. The transport dispatches to the same registered tool schemas.
  4. Valid input invokes the graph; invalid input returns a tool error instead.
  5. MemoryGraph returns a result or error to the tool handler.
  6. The SDK hands the MCP response back to the selected transport.
  7. The agent interprets the result and checks current files itself.

The documented Worker implementation has no caller authentication. Managed read transactions and response limits do not make agent-authored Cypher safe for untrusted callers. A deployment with access to a sensitive graph requires access controls and restricted database credentials.

search accepts a natural-language problem description. OpenRouter embeds the query with voyageai/voyage-4. Neo4j’s separately created attempt_embedding index returns up to 200 similar attempts. Mentis groups the results by (repository, taskId), keeps at most five attempts per task, and uses typesafe/jev-1.13 to score usefulness. If any Jev score fails, the whole search falls back to vector similarity. Embedding and vector-index failures return tool errors.

The result is a list of tasks, each with one matched attempt preview. Similarity and usefulness scores do not indicate correctness. Relevant attempts outside the vector result pool may be omitted. Returned attempts may include failed checks or outdated conclusions.

Read task history

recall runs agent-supplied Cypher to read attempts directly, without OpenRouter or a vector index. Queries should normally use both repository and task ID as parameters. The tool returns columns and rows as JSON text in the MCP result. It flags results truncated at 100 rows or 512,000 serialized bytes, and has a five-second read timeout. These limits do not bound the cost of every possible query.

How does exact-history recall differ from semantic search?
Step 1 of 6 The agent asks for exact history using parameterized Cypher, not a natural-language prompt.
Read all steps
  1. The agent asks for exact history using parameterized Cypher, not a natural-language prompt.
  2. The tool validates Cypher and rejects reserved parameter keys.
  3. Invalid input returns a tool error with no database query.
  4. On a separate valid request, Neo4j executes a bounded read transaction without OpenRouter.
  5. The response retains at most 100 rows and 512,000 serialized bytes; truncation is flagged.
  6. Recall returns JSON text for the agent to interpret against its current checkout.

Include failed attempts and correction fields in the Cypher projection when reviewing task history. Verify relevant observations against the current checkout before using them.

Record an attempt

The agent calls record_attempt with the repository, task ID, code context, action, affected files, and observed result. An optional inference is stored separately from the observation. The agent should report a check as passed or failed only if it ran the check.

Mentis embeds the document text before starting the Neo4j write transaction. If embedding fails or does not return exactly 1,024 finite numbers, nothing is written. On success, one transaction merges the repository and task and adds a new attempt with a UUID. Subsequent calls add attempts rather than replacing earlier records.

Why can a failed embedding never leave a half-written attempt?
Step 1 of 7 The agent records a real action and observation. Missing check means unverified.
Read all steps
  1. The agent records a real action and observation. Missing check means unverified.
  2. Validation requires nonempty fields and affected files.
  3. If input is invalid, the call fails without embedding or writing.
  4. On a new valid request, OpenRouter creates a document embedding before the transaction.
  5. If embedding fails or is invalid, no repository, task or attempt is written.
  6. On another valid call, the vector succeeds and one transaction merges the repository and task.
  7. The transaction creates a new attempt. Repeating the call appends another.

Mark an outdated conclusion

mark_conclusion_outdated matches an attempt by repository and attempt ID. It stores a reason and timestamp and can also store an agent-reported latest commit. The original action, observation, check result, and embedding remain unchanged. Search can still return the attempt with its correction. Recall queries must explicitly request the correction properties.

Calling the tool again replaces the previous reason and timestamp; it does not preserve a correction history. New evidence should be stored as a new attempt. forget_attempt deletes an attempt from the live graph and is separate from marking a conclusion outdated.

What changes when a conclusion becomes outdated?
Step 1 of 7 An old attempt contains observations and an inference, not a permanent truth.
Read all steps
  1. An old attempt contains observations and an inference, not a permanent truth.
  2. The agent checks current code and decides the inference is outdated.
  3. The correction matches one attempt by repository identity and UUID.
  4. The original action, observation, check, and embedding are preserved.
  5. A write sets outdatedReason and outdatedAt; repeating it overwrites these fields.
  6. Search can still match the original embedding and shows the outdated marker.
  7. Recall can read both the original evidence and selected correction fields.

Data handling and verification

Query text and attempt content can be sent to OpenRouter for embedding or scoring. Do not include secrets. Git commit and dirty-state fields are supplied by the agent and are not verified by Mentis. Before using a retrieved attempt, inspect its history and check its observations against the current files.