CapyNodes

An interactive system design canvas where an AI Judge scores your architecture for scalability, reliability, and failure modes.

  • Python
  • Langchain
  • Groq
  • Django
  • PostgreSQL
  • LLMs
impact:
  • 50+ Active users
  • 250+ Evaluation sessions
  • 3-stage Eval pipeline

CapyNodes is a system-design practice environment built around an interactive architecture canvas and an AI judge. Instead of treating a diagram as a static answer, it represents a design as structured graph state, evaluates it against problem constraints, and returns a scored explanation of the system’s trade-offs and failure modes.

The central problem is not drawing boxes and arrows. It is turning an ambiguous system-design answer into feedback that is fast enough to be useful, deterministic where it should be, and nuanced where rules alone are insufficient. CapyNodes addresses this with a real-time collaboration layer and a hybrid evaluation pipeline.


Demo


System architecture

The application separates interactive diagram state from the evaluation workload:

React Flow canvas
      |
      | HTTP + authenticated WebSocket messages
      v
Django + Django REST Framework
      |
      +-- Django Channels + Redis channel layer --> collaboration participants
      |
      +-- PostgreSQL JSONB --> questions, submissions, evaluations, diagrams
      |
      +-- evaluation engine --> rules -> LLM -> score aggregation

The frontend is a Next.js application using React Flow for the canvas. The backend is Django with Django REST Framework for the application API and Django Channels for bidirectional collaboration. Redis provides the channel layer used to broadcast changes between connected clients. PostgreSQL persists the durable state; the project uses JSONB for diagrams and LLM evaluation results because both are structured but evolve more quickly than a fixed relational schema.

This division keeps the latency-sensitive interaction path separate from the slower model-backed evaluation path. A cursor move should not wait for a database write or an LLM response. An evaluation, conversely, can work from a durable graph snapshot rather than transient browser state.


Diagram data model

A system design is a graph of nodes, edges, and UI metadata. The database persists the graph as JSONB in a submission or collaboration session, allowing the canvas to retain its native React Flow shape without a join-heavy node and edge schema.

The core records are:

  • Question stores a design prompt, its constraints, and an optional ideal solution as structured JSON.
  • Submission stores a participant’s diagram, evaluation result, and score.
  • CollaborationSession stores the latest shared diagram state for a room.
  • LLMCallLog records prompt tokens, completion tokens, and latency for model observability.

The evaluation engine does not score raw canvas geometry. It normalizes the graph first, removing UI-only properties such as node position and styling, and serializing the remaining structural representation deterministically. A SHA-256 hash of that normalized form identifies semantically identical diagrams. This creates a cache key that is stable when users only rearrange the canvas and allows a prior evaluation to be reused when the graph, prompt version, and model version have not changed.


Real-time collaboration

Collaboration is implemented with authenticated Django Channels WebSockets. A participant joins a room named collab_<sessionId>, and the CollaborationConsumer relays two distinct classes of event.

state_update carries a diagram change. The backend broadcasts the update to the other participants and persists the latest state asynchronously through database_sync_to_async, keeping ORM work off the asynchronous consumer path. Persisting graph state makes a collaboration session recoverable after clients reconnect.

cursor_update carries only a participant’s pointer coordinates. It is broadcast but not stored. Cursor traffic is high frequency, ephemeral, and irrelevant to reconstructing the diagram, so persisting it would add database load without increasing durability.

The consumer filters a sender’s own broadcast by channel name. This prevents a browser from re-applying its local edit after the server has fanned it out, avoiding feedback loops and duplicate canvas updates. The resulting consistency model is pragmatic: participants see low-latency broadcasts, while the database converges on the most recently persisted graph state.


Hybrid evaluation pipeline

The AI judge runs three stages, each responsible for a different kind of confidence.

1. Structural validation

The rule engine evaluates graph properties that should not depend on a generative model. It detects empty or disconnected designs, orphaned components, invalid connections, missing flows, and recognized anti-patterns such as a single point of failure. It also evaluates explicit problem constraints, for example whether a high-throughput design has the expected buffering or storage components.

This stage is deterministic and fast. Its output is a structural score plus critical issues. It gives the product a reliable floor: basic graph errors do not need expensive probabilistic reasoning.

2. LLM architecture review

The LLM evaluator sends the normalized diagram and the design prompt to Qwen 32B through Groq. The prompt makes the model first explain the architecture it inferred, then analyze strengths and weaknesses, and finally score the design across relevant dimensions such as scalability, reliability, performance, and security.

This stage handles the part rules cannot encode economically: whether a cache placement is justified, whether a data model suits the access pattern, or whether two individually valid components make sense together. The model returns both scores and reasoning so feedback is actionable rather than a single opaque number.

3. Aggregation and calibration

The score aggregator combines the deterministic and model-backed outputs into one dimensional scorecard. The rule-based signal accounts for roughly 30% of the result and LLM judgment roughly 70%; difficulty multipliers and normalization are then applied to keep scores comparable across question levels.

If the LLM is unavailable, a fallback evaluator compares the submission with the question’s ideal solution using node types, connection patterns, and architecture patterns. It produces a basic score and explicitly indicates that the model-backed review was unavailable rather than presenting fallback output as an equivalent evaluation.


Evaluation quality and operations

A scoring system needs measurement beyond whether a request returned successfully. CapyNodes records token usage and latency per LLM call, then surfaces quality and operational signals in a Streamlit dashboard. The dashboard tracks evaluation quality trends, p50 and p95 model latency, token cost, and failures such as provider or decoding errors.

A curated Golden Set of diagrams provides a regression suite for prompt and model changes. Each case has an expected evaluation baseline; a candidate change is compared against it before becoming the new judge behavior. A secondary Gemini model runs offline cross-checks against the primary Qwen evaluator to identify hallucinations and score variance without placing a second model in the interactive request path.

The project therefore treats model output as a versioned dependency with measurable behavior, not as an infallible evaluator. Graph normalization and caching reduce repeat work, deterministic rules catch obvious defects, golden cases protect against prompt regressions, and offline cross-checks investigate disagreement.


Design limits

The collaboration protocol prioritizes low-latency shared editing over a full operational-transform or CRDT implementation. Diagram updates are broadcast as state snapshots and cursors are intentionally ephemeral. The judge is also a feedback system, not a formal verifier: it combines explicit structural checks with model reasoning and exposes the reasoning so a user can challenge the result.

That scope keeps the product focused on its loop: draw an architecture, receive technically grounded feedback, revise the graph, and learn from the next evaluation.

Site search and portfolio assistant