Skip to content

About

Interactive code assessment playground: submit a solution and three Gemini agents grade it against rubric-grounded RAG retrieval, then coach you Socratically.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DevSkill Arena

An interactive code assessment playground. Solve a challenge in a Monaco editor, submit it, and three Gemini agents grade it against retrieved skill rubrics — then coach you toward the fix without handing it over.

Runs entirely on its own: no database, no backend service, no infrastructure. One npm install, one npm run dev.

Stack: Next.js 14 (App Router) · TypeScript · Tailwind CSS · Monaco Editor · Framer Motion · GSAP · Google Gemini


Setup

npm install
cp .env.example .env.local     # add your Gemini API key
npm run dev                    # http://localhost:3000

Browsing challenges works with no key at all. The Evaluate button needs a Gemini key, because grading makes real LLM calls — get one free at aistudio.google.com/apikey and put it in .env.local:

GEMINI_API_KEY=your-key-here

If the key is missing, POST /api/assessments/evaluate returns 503 MISSING_API_KEY with an explanatory message. It never invents a score.

Commands

Command What it does
npm run dev Dev server on :3000
npm run build Production build
npm start Serve the production build
npm run typecheck tsc --noEmit
npm run embed Regenerate data/embeddings.json — only needed if you edit data/skills.json

Security note

GEMINI_API_KEY has no NEXT_PUBLIC_ prefix and is read only inside server-only modules (lib/server/*, lib/retrieval/*, guarded by import 'server-only'). It is never bundled for the browser and never sent to the client. Verified against the production build: the client bundle contains no key, no rubric text, no agent prompts, and no embedding vectors.


How evaluation works

POST /api/assessments/evaluate runs the whole pipeline inside a single Next.js route handler:

submitted code
  │
  ├─ 1. embed the code            gemini-embedding-001, 768-dim, RETRIEVAL_QUERY
  ├─ 2. retrieveContext()         cosine vs data/embeddings.json, > 0.35, top 4
  ├─ 3. Agent 1  Syntax & Spec    scores rubric compliance      0-100
  ├─ 4. Agent 2  Test & Logic     one test per numbered constraint
  ├─ 5. Agent 3  Socratic Mentor  questions, never answers
  └─ 6. score = 0.4 × syntax + 0.6 × logic,  pass at ≥ 70

The three agent system prompts in lib/server/gemini.ts are the assessment product. They are calibrated against the scoring weights; changing their wording changes every score the app produces.

Graceful degradation. Any single step failing sets degraded: true and appends to warnings[] rather than failing the request — partial feedback is worth more to a learner than an error page, and the feedback drawer renders that state. Only if both scoring agents fail does the request return 502.

Retrieval

Retrieval sits behind one interface, Retriever.retrieveContext(), declared in lib/retrieval/types.ts. The pipeline, the agents and the route handler depend on nothing else, so swapping strategy — keyword, hybrid, a real vector database — is a new file plus one line in lib/retrieval/index.ts.

The current implementation (lib/retrieval/embedding.ts) is vector search without a vector database:

Model gemini-embedding-001
Width 768 (Matryoshka truncation)
Normalization L2 — required at any width ≠ 3072, which is the only pre-normalized one
Corpus task type RETRIEVAL_DOCUMENT
Query task type RETRIEVAL_QUERY
Metric cosine
Threshold 0.35, strict >
top-k 4
Chunking none — one rubric, one vector
Document text `${name} (${category})\n${rubric_context}`
Storage data/embeddings.json, read with fs at runtime

Corpus vectors are precomputed by npm run embed and committed, so the app runs for anyone who clones it without an embedding-time key. Only the per-request query embedding needs the API key. The file is read from disk rather than statically imported, so the vectors can never be pulled into a client bundle.

The corpus is rubrics only. Challenge text is never embedded — it reaches the agents as plain prompt text.


Editing challenges

Challenges live in data/challenges.json — a plain array. Add an object, save, reload. No embedding step, no rebuild in dev.

{
  "id": "a-unique-uuid",
  "title": "Responsive Card Row with Flexbox",
  "description": "Build a horizontal row of 3 cards that: (1) uses display:flex on .card-row, (2) ... Constraints: no floats, no !important.",
  "starter_code": "<style>\n.card-row {\n  /* your layout here */\n}\n</style>",
  "language": "css",
  "difficulty": "easy",
  "created_at": "2026-02-02T09:00:00Z"
}
Field Notes
id Any unique string. Appears in the URL as /playground?challenge=<id>.
title Card heading and playground header.
description The brief. Number your constraints — Agent 2 produces one test per numbered constraint, so (1) … (2) … directly shapes the grading.
starter_code Pre-filled editor contents. Remember to escape newlines as \n in JSON.
language Drives Monaco syntax highlighting and the live preview. css (an HTML doc with a <style> block) gets the Code/Preview tabs; go, javascript, typescript get the editor only.
difficulty Exactly easy, medium or hard — these map to the coloured badges.
created_at ISO 8601. The catalog sorts newest first.

The language filter pills are derived from whatever language values are present, so a new language appears automatically.

Editing rubrics

Grading rubrics live in data/skills.json, and their embeddings are cached in data/embeddings.json.

If you edit, add or remove a rubric, re-run npm run embed and commit the result:

npm run embed

You do not need to re-run it after editing data/challenges.json. If a rubric changes without re-embedding, the app still works but adds a degraded warning naming the stale rubric, surfaced in the feedback drawer — the corpus loader hashes each document and compares.


Downloading a report

The feedback drawer has a Download PDF button. It calls window.print() and the browser's own PDF engine does the rest — choose "Save as PDF" as the destination in the print dialog.

There is no PDF library and no server round trip. The report renders from the evaluation result already in client state, so downloading costs nothing, spends no API quota, re-runs no agents, and works offline.

Section Source
Header + score summary Challenge metadata, final score, the 40/60 breakdown, pass mark
Points to improve Derived: high-severity violations → failed constraints → medium → low
Questions to guide you Agent 3 hints, next step, encouragement
What went well Agent 1 strengths
Constraint results Agent 2, every test with expected vs actual
Rubrics used Retrieved rubrics with cosine similarity
Appendix The submitted code, on its own page

"Points to improve" is a pure function (buildImprovementPlan in lib/report.ts) over the agent output — most actionable first. Agent 3's Socratic hints are deliberately kept in a separate section rather than mixed in, since a question is not a defect.

Print implementation notes

  • The printable report (components/EvaluationReport.tsx) is portalled to document.body. The drawer is position: fixed and carries a Framer Motion transform, and fixed or transformed ancestors print unreliably — usually only the first page. As a direct child of <body> the report is a normal static block and paginates correctly.
  • @media print in globals.css hides every other child of <body> and reveals #evaluation-print-root.
  • No reliance on background colours. Browsers print with "background graphics" off by default, so hierarchy comes from rules, borders and weight. Severity and pass/fail are text tags, not coloured chips.
  • break-inside: avoid is applied to item-level blocks, not whole sections. A section taller than one page cannot honour it, and browsers respond by ejecting a blank page. Keeping each violation, hint and test atomic is what actually prevents a finding being split mid-thought; break-after: avoid on headings stops a heading stranding at a page foot.
  • Margins are set twice, on purpose. @page { margin } sits at the top level of globals.css rather than inside @media print, because some engines ignore a nested @page. It handles the top and bottom of middle pages, which container padding cannot — block padding applies only to the first and last page. .report then carries its own horizontal padding as a floor, so if a user's print dialog has Margins: None the content still is not flush against the paper edge. On A4 the text column ends up ~166mm wide with @page honoured, ~190mm without.
  • document.title is swapped to DevSkill Arena - <challenge> - Score <n> for the duration of the print, because that is what browsers use for the suggested filename, then restored on afterprint (with a timeout fallback for browsers that never fire it).

Project layout

app/
  layout.tsx  page.tsx  globals.css        challenge catalog + theme
  playground/page.tsx                      editor, preview, evaluate
  api/assessments/evaluate/route.ts        the pipeline entry point
components/
  CodeEditor       Monaco wrapper, dark theme
  LivePreview      sandboxed iframe for CSS/HTML challenges
  ScoreMeter       GSAP score dial
  AgentFeedback    results drawer + Download PDF
  EvaluationReport print-only report
lib/
  types.ts         domain + wire types
  api.ts           data access (local JSON + evaluate call)
  report.ts        pure report derivation
  server/          config  gemini  pipeline  challenges     server-only
  retrieval/       types  document  embedding  index        server-only
data/
  challenges.json  hand-edited
  skills.json      hand-edited
  embeddings.json  generated by `npm run embed`, committed
scripts/embed.mjs  zero-dependency embedding script

Design notes and limitations

  • Nothing is persisted. There is no database, so no attempt history survives a reload and submission_id comes back empty (the drawer shows unsaved). Download the PDF to keep a record. The "vs last attempt" delta on the score dial works in-memory for the session.
  • GEMINI_MAX_OUTPUT_TOKENS defaults to 8192, and needs to be generous. gemini-3.6-flash charges ~1400–1700 reasoning tokens against the same budget as the output, so a low ceiling starves Agent 2 — the agent with the longest prompt — which then returns truncated or empty JSON. Since Agent 2 is 60% of the score, that degrades the result quietly. Anything below ~4096 is risky.
  • Agent calls declare a responseSchema. responseMimeType: application/json alone is not sufficient: Agent 2 intermittently emitted unquoted string values, which fails JSON.parse. Declaring the schema switches the model into constrained decoding.
  • The corpus is demo-sized. Five rubrics means top-k of 4 does very little discriminating — most inputs retrieve nearly everything above the threshold. Retrieval ranks the right rubric first, but a realistic corpus is needed before the threshold and top-k values are meaningfully tuned.
  • Model churn is a real cost. The embedding width, chat model and token budget are all environment-configurable precisely because Gemini model lifecycles are short. Pin them deliberately.
  • Code is analysed, not executed. Agent 2 reasons about behaviour analytically; there is no sandboxed runner and no real test execution. The live preview renders CSS/HTML submissions in an iframe with sandbox="allow-scripts" and no allow-same-origin, so submitted code cannot reach the app's DOM, cookies or storage.

License

MIT — see LICENSE.

About

Interactive code assessment playground: submit a solution and three Gemini agents grade it against rubric-grounded RAG retrieval, then coach you Socratically.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages