An interactive code assessment playground. Solve a challenge in a Monaco editor, submit it, and three Gemini agents grade it against retrieved skill rubrics — then coach you toward the fix without handing it over.
Runs entirely on its own: no database, no backend service, no infrastructure.
One npm install, one npm run dev.
Stack: Next.js 14 (App Router) · TypeScript · Tailwind CSS · Monaco Editor · Framer Motion · GSAP · Google Gemini
npm install
cp .env.example .env.local # add your Gemini API key
npm run dev # http://localhost:3000Browsing challenges works with no key at all. The Evaluate button needs a
Gemini key, because grading makes real LLM calls — get one free at
aistudio.google.com/apikey and put it in
.env.local:
GEMINI_API_KEY=your-key-here
If the key is missing, POST /api/assessments/evaluate returns 503
MISSING_API_KEY with an explanatory message. It never invents a score.
| Command | What it does |
|---|---|
npm run dev |
Dev server on :3000 |
npm run build |
Production build |
npm start |
Serve the production build |
npm run typecheck |
tsc --noEmit |
npm run embed |
Regenerate data/embeddings.json — only needed if you edit data/skills.json |
GEMINI_API_KEY has no NEXT_PUBLIC_ prefix and is read only inside
server-only modules (lib/server/*, lib/retrieval/*, guarded by
import 'server-only'). It is never bundled for the browser and never sent to
the client. Verified against the production build: the client bundle contains
no key, no rubric text, no agent prompts, and no embedding vectors.
POST /api/assessments/evaluate runs the whole pipeline inside a single Next.js
route handler:
submitted code
│
├─ 1. embed the code gemini-embedding-001, 768-dim, RETRIEVAL_QUERY
├─ 2. retrieveContext() cosine vs data/embeddings.json, > 0.35, top 4
├─ 3. Agent 1 Syntax & Spec scores rubric compliance 0-100
├─ 4. Agent 2 Test & Logic one test per numbered constraint
├─ 5. Agent 3 Socratic Mentor questions, never answers
└─ 6. score = 0.4 × syntax + 0.6 × logic, pass at ≥ 70
The three agent system prompts in lib/server/gemini.ts are the assessment
product. They are calibrated against the scoring weights; changing their
wording changes every score the app produces.
Graceful degradation. Any single step failing sets degraded: true and
appends to warnings[] rather than failing the request — partial feedback is
worth more to a learner than an error page, and the feedback drawer renders
that state. Only if both scoring agents fail does the request return 502.
Retrieval sits behind one interface, Retriever.retrieveContext(), declared in
lib/retrieval/types.ts. The pipeline, the agents and the route handler depend
on nothing else, so swapping strategy — keyword, hybrid, a real vector database
— is a new file plus one line in lib/retrieval/index.ts.
The current implementation (lib/retrieval/embedding.ts) is vector search
without a vector database:
| Model | gemini-embedding-001 |
| Width | 768 (Matryoshka truncation) |
| Normalization | L2 — required at any width ≠ 3072, which is the only pre-normalized one |
| Corpus task type | RETRIEVAL_DOCUMENT |
| Query task type | RETRIEVAL_QUERY |
| Metric | cosine |
| Threshold | 0.35, strict > |
| top-k | 4 |
| Chunking | none — one rubric, one vector |
| Document text | `${name} (${category})\n${rubric_context}` |
| Storage | data/embeddings.json, read with fs at runtime |
Corpus vectors are precomputed by npm run embed and committed, so the app
runs for anyone who clones it without an embedding-time key. Only the
per-request query embedding needs the API key. The file is read from disk rather
than statically imported, so the vectors can never be pulled into a client
bundle.
The corpus is rubrics only. Challenge text is never embedded — it reaches the agents as plain prompt text.
Challenges live in data/challenges.json — a plain array. Add an object,
save, reload. No embedding step, no rebuild in dev.
{
"id": "a-unique-uuid",
"title": "Responsive Card Row with Flexbox",
"description": "Build a horizontal row of 3 cards that: (1) uses display:flex on .card-row, (2) ... Constraints: no floats, no !important.",
"starter_code": "<style>\n.card-row {\n /* your layout here */\n}\n</style>",
"language": "css",
"difficulty": "easy",
"created_at": "2026-02-02T09:00:00Z"
}| Field | Notes |
|---|---|
id |
Any unique string. Appears in the URL as /playground?challenge=<id>. |
title |
Card heading and playground header. |
description |
The brief. Number your constraints — Agent 2 produces one test per numbered constraint, so (1) … (2) … directly shapes the grading. |
starter_code |
Pre-filled editor contents. Remember to escape newlines as \n in JSON. |
language |
Drives Monaco syntax highlighting and the live preview. css (an HTML doc with a <style> block) gets the Code/Preview tabs; go, javascript, typescript get the editor only. |
difficulty |
Exactly easy, medium or hard — these map to the coloured badges. |
created_at |
ISO 8601. The catalog sorts newest first. |
The language filter pills are derived from whatever language values are
present, so a new language appears automatically.
Grading rubrics live in data/skills.json, and their embeddings are cached
in data/embeddings.json.
If you edit, add or remove a rubric, re-run npm run embed and commit the
result:
npm run embedYou do not need to re-run it after editing data/challenges.json. If a
rubric changes without re-embedding, the app still works but adds a degraded
warning naming the stale rubric, surfaced in the feedback drawer — the corpus
loader hashes each document and compares.
The feedback drawer has a Download PDF button. It calls window.print()
and the browser's own PDF engine does the rest — choose "Save as PDF" as the
destination in the print dialog.
There is no PDF library and no server round trip. The report renders from the evaluation result already in client state, so downloading costs nothing, spends no API quota, re-runs no agents, and works offline.
| Section | Source |
|---|---|
| Header + score summary | Challenge metadata, final score, the 40/60 breakdown, pass mark |
| Points to improve | Derived: high-severity violations → failed constraints → medium → low |
| Questions to guide you | Agent 3 hints, next step, encouragement |
| What went well | Agent 1 strengths |
| Constraint results | Agent 2, every test with expected vs actual |
| Rubrics used | Retrieved rubrics with cosine similarity |
| Appendix | The submitted code, on its own page |
"Points to improve" is a pure function (buildImprovementPlan in
lib/report.ts) over the agent output — most actionable first. Agent 3's
Socratic hints are deliberately kept in a separate section rather than mixed in,
since a question is not a defect.
- The printable report (
components/EvaluationReport.tsx) is portalled todocument.body. The drawer isposition: fixedand carries a Framer Motion transform, and fixed or transformed ancestors print unreliably — usually only the first page. As a direct child of<body>the report is a normal static block and paginates correctly. @media printinglobals.csshides every other child of<body>and reveals#evaluation-print-root.- No reliance on background colours. Browsers print with "background graphics" off by default, so hierarchy comes from rules, borders and weight. Severity and pass/fail are text tags, not coloured chips.
break-inside: avoidis applied to item-level blocks, not whole sections. A section taller than one page cannot honour it, and browsers respond by ejecting a blank page. Keeping each violation, hint and test atomic is what actually prevents a finding being split mid-thought;break-after: avoidon headings stops a heading stranding at a page foot.- Margins are set twice, on purpose.
@page { margin }sits at the top level ofglobals.cssrather than inside@media print, because some engines ignore a nested@page. It handles the top and bottom of middle pages, which container padding cannot — block padding applies only to the first and last page..reportthen carries its own horizontal padding as a floor, so if a user's print dialog has Margins: None the content still is not flush against the paper edge. On A4 the text column ends up ~166mm wide with@pagehonoured, ~190mm without. document.titleis swapped toDevSkill Arena - <challenge> - Score <n>for the duration of the print, because that is what browsers use for the suggested filename, then restored onafterprint(with a timeout fallback for browsers that never fire it).
app/
layout.tsx page.tsx globals.css challenge catalog + theme
playground/page.tsx editor, preview, evaluate
api/assessments/evaluate/route.ts the pipeline entry point
components/
CodeEditor Monaco wrapper, dark theme
LivePreview sandboxed iframe for CSS/HTML challenges
ScoreMeter GSAP score dial
AgentFeedback results drawer + Download PDF
EvaluationReport print-only report
lib/
types.ts domain + wire types
api.ts data access (local JSON + evaluate call)
report.ts pure report derivation
server/ config gemini pipeline challenges server-only
retrieval/ types document embedding index server-only
data/
challenges.json hand-edited
skills.json hand-edited
embeddings.json generated by `npm run embed`, committed
scripts/embed.mjs zero-dependency embedding script
- Nothing is persisted. There is no database, so no attempt history
survives a reload and
submission_idcomes back empty (the drawer showsunsaved). Download the PDF to keep a record. The "vs last attempt" delta on the score dial works in-memory for the session. GEMINI_MAX_OUTPUT_TOKENSdefaults to 8192, and needs to be generous.gemini-3.6-flashcharges ~1400–1700 reasoning tokens against the same budget as the output, so a low ceiling starves Agent 2 — the agent with the longest prompt — which then returns truncated or empty JSON. Since Agent 2 is 60% of the score, that degrades the result quietly. Anything below ~4096 is risky.- Agent calls declare a
responseSchema.responseMimeType: application/jsonalone is not sufficient: Agent 2 intermittently emitted unquoted string values, which failsJSON.parse. Declaring the schema switches the model into constrained decoding. - The corpus is demo-sized. Five rubrics means top-k of 4 does very little discriminating — most inputs retrieve nearly everything above the threshold. Retrieval ranks the right rubric first, but a realistic corpus is needed before the threshold and top-k values are meaningfully tuned.
- Model churn is a real cost. The embedding width, chat model and token budget are all environment-configurable precisely because Gemini model lifecycles are short. Pin them deliberately.
- Code is analysed, not executed. Agent 2 reasons about behaviour
analytically; there is no sandboxed runner and no real test execution. The
live preview renders CSS/HTML submissions in an iframe with
sandbox="allow-scripts"and noallow-same-origin, so submitted code cannot reach the app's DOM, cookies or storage.
MIT — see LICENSE.