Skip to content

About

CareerSphere — grounded career exploration on Singapore's official SkillsFuture and live jobs data (PyCon SG 2026 hackathon).

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

CareerSphere

CareerSphere is a career exploration app built for the PyCon SG 2026 Jobs & Skills hackathon.

It tries to answer three practical questions:

  1. Where do I stand right now?
  2. Where can I realistically go next?
  3. What should I do to get there?

A user can describe their background, goals, constraints, salary expectations, timeline, or general uncertainty in plain English. The app then tries to give a reasonable grounded answer using SkillsFuture and jobs data.

The basic idea is simple: AI can help understand what the user is asking, but the important facts should still come from data and deterministic code.

What It Does

CareerSphere takes a user profile or resume and turns it into a guided career analysis:

  1. Extract skills and evidence from the user's text.
  2. Match those skills to official SkillsFuture skills.
  3. Compare the user against official role requirements.
  4. Show a 3D sphere of role compatibility.
  5. Explain what roles are realistic now, what is further away, and why.
  6. Suggest a practical next move, course, or project.
  7. Show matching jobs from MyCareersFuture-style data.

The app is intentionally more open than a fixed form. A user can ask things like:

  • "I want a higher salary but only have 6 months."
  • "I want to switch into data but I don't want to become a software engineer."
  • "Which nearby jobs can I apply for now?"
  • "What is a realistic move from my current background?"

When the goal is too far, the timeline is too short, or the salary expectation looks unrealistic, the app should say that plainly instead of pretending everything is equally achievable.

Architecture

careersphere system architecture

An LLM agent orchestrates the analysis by making grounded tool calls against real data: the SkillsFuture Skills Framework (baked into an in-process DuckDB snapshot), live jobs from MyCareersFuture, courses from the SkillsFuture directory, and OpenAI for parsing and semantic matching. The important facts come from the data and deterministic tools, not the model.

UX Shape

The app uses a mostly linear flow: onboarding, analysis, sphere, then recommendations.

That was a deliberate UX choice. Career advice tools can get messy quickly because there are too many possible directions, and a dashboard can make the user do a lot of the stitching themselves. So instead of building a space where people have to jump between tabs, we tried to make the product move in one direction: first understand the user, then show the landscape, then explain what is realistic, then suggest what to do next.

That same thinking affected the demo format. We deprioritized a separate video walkthrough because the product flow needed to explain itself. Instead, we added text around the key features so the user can understand what they are seeing while moving through the app, rather than needing a separate video to decode the interface.

That is also why the sphere is not meant to stand alone. A 3D view of many roles is interesting, but by itself it can also be confusing, especially when there are a lot of jobs and role directions to scan. The sphere gives the map feeling; the surrounding explanations, ranked roles, gaps, jobs, and evidence explain what the map actually means.

Why A Sphere?

Partly because we thought it was cool.

It also made sense for the kind of matching problem we were working with. A lot of the underlying work is about similarity: matching a user profile to skills, roles, and jobs. Embeddings and cosine similarity already treat text as points in a high-dimensional space.

The 3D sphere is obviously a simplification. It flattens a lot of information. But as a visual metaphor, it lines up with the math better than a plain list does: nearby roles feel closer, distant roles feel further away, and clusters give a rough sense of related directions.

Still, the sphere is not the whole product. It is a way into the analysis, not the analysis itself.

Datasets

Dataset Why we used it
SkillsFuture Skills Framework Official roles, role descriptions, required skills, and proficiency levels.
SkillsFuture Unique Skills List Canonical skill vocabulary and CASL / Emerging Skills signals.
TSC-to-Unique Skills Mapping Connects sector-specific framework skills to canonical skills.
SkillsFuture Course Directory Lets us suggest real courses for skill gaps. We also tried using the live course API path, which should be maintained, but some course-detail links still return 404s in practice.
MyCareersFuture job postings Gives market-facing job signals and listed skills.
Apify experiments Used for testing fresh web/profile evidence, though scraping was not reliable enough to make core.

Sources:

Raw SkillsFuture files, DuckDB files, generated embedding caches, and local secrets are not committed.

Responsible AI And Interpretability

We tried to keep a hard boundary between language work and factual computation.

The app stays mostly inside the OpenAI ecosystem for AI features:

  • Structured Outputs for profile parsing, so resume/profile text becomes predictable structured data instead of loose JSON;
  • embeddings for semantic matching between user skills, SkillsFuture skills, and role/job text;
  • tool-calling orchestration over allowlisted deterministic backend functions;
  • plain-English explanations from computed JSON.

The LLM can:

  • parse messy profile text into structured skills and evidence;
  • infer conservative skill levels from evidence;
  • route between allowlisted backend tools;
  • explain the computed result in plain English.

We also chose semantic embeddings because matching against only the role and skill text in the raw SkillsFuture tables was too brittle. Some requirements in the base data can be surprising in context, such as network-related skills showing up for a UX designer role, and exact keyword matching tended to overreact to those oddities. Embeddings gave us a better way to compare user evidence, skill descriptions, role descriptions, and job text semantically, while still keeping deterministic thresholds and fallbacks around the final match.

The LLM cannot:

  • create new roles;
  • change SkillsFuture skill definitions;
  • invent proficiency requirements;
  • decide fit scores;
  • rank jobs without evidence;
  • invent salaries, courses, or source rows.

The scoring path is deterministic:

user profile
  -> OpenAI Structured Outputs profile parse
  -> OpenAI embeddings / deterministic fallback for skill matching
  -> role requirement lookup
  -> role fit scoring
  -> gap ranking
  -> course lookup
  -> job ranking
  -> explanation

The key formulas and decisions live in code and docs, not in prompts. For example, Framework Fit is computed from official role-skill-proficiency rows, and job matching uses listed-skill coverage rather than guessed proficiency levels.

Interpretability matters because a career recommendation is only useful if the user can inspect it. The app should be able to answer: "Which row or skill caused this recommendation?"

Relevant docs:

Tech Stack

Python was the workhorse here.

We were not trying to use Python in some especially new or frontier way. We used it because it is good for exactly this kind of project: ingest data, clean it, glue systems together, write scoring logic, expose an API, and test quickly.

  • Python 3.11
  • FastAPI
  • DuckDB / MotherDuck
  • Pydantic
  • pandas / pyarrow / numpy
  • pytest
  • OpenAI Responses API
  • OpenAI Structured Outputs
  • OpenAI embeddings
  • React + Vite + TypeScript
  • react-three-fiber / Three.js
  • Google Cloud Run
  • Google Cloud Scheduler
  • Apify experiments

OpenAI plus Python was our main iteration loop. Python gave us a stable backend for deterministic logic, while OpenAI helped with language understanding, embeddings, and agentic orchestration.

How We Built It

A large part of the technical work was done with AI assistance. That was intentional.

The aim was to see how much a small team could ship by leaning into prompting, agentic coding, and fast review loops. We split across Codex CLI and the Codex app depending on the task, with the CLI being useful for longer backend loops and the app being useful for repo-aware iteration and UI work. The timing was also very hackathon-like: one of us had just landed back in Singapore from exchange right before PyCon, so some of the work happened in a slightly compressed, jet-lagged state.

AI tools let us focus more time on design direction, product tradeoffs, and iteration speed. They also generated a lot of the actual code. But that did not remove the need for technical understanding. If anything, it made it more obvious.

When iteration gets too rapid with AI, technical debt appears quickly. The repo is honest about that. Features can land fast, but someone still needs to understand the architecture, inspect generated code, test behavior, catch wrong assumptions, and decide what is worth keeping.

One of the most time-consuming parts was repeatedly testing the agent flow itself. The happy path looked simple on paper - parse profile, match skills, find roles, rank gaps, explain - but small changes in prompts, tool order, fallback behavior, or input wording could change the final result quite a bit. We spent a lot of time running the app, reading traces, adjusting prompts and guardrails, and checking whether the recommendations actually felt reasonable.

The deterministic side needed tuning too. We had to try different weights for things like career-stage mismatch, how much to penalize a role that is technically related but unrealistic, how strongly to fold in the user's stated intent, and how much the system should prefer close realistic moves over more ambitious ones. Those numbers were not left to the LLM, but they still needed human judgment and repeated testing.

In the current scoring, "fit" and "what we serve next" are related but not the same thing. The app first asks whether the role is semantically relevant to the user's background, then uses official SkillsFuture requirement coverage as a grounded lift and audit trail. That coverage matters a lot for evidence, gaps, and readiness, but it does not completely dominate the headline match because some framework rows are broad or surprising in context.

After that, the app applies practical steering. Career-stage mismatch can demote roles that are technically related but unrealistic, and the user's stated intent can tilt which reachable destination or next action is emphasized. The important part is that intent only steers among credible options; it does not rewrite where the user fits. Gap ranking follows the same idea: prefer skills that are close enough to build, reused across roles, visible in jobs, and supported by official future-skill signals.

So AI was not a replacement for technicality. It was an accelerant. Useful, sometimes messy, and only really effective when paired with enough engineering judgment to know when to trust it and when to push back.

Our working loop was roughly:

discuss idea
  -> ask AI to critique or implement
  -> test it in the app
  -> argue about whether it makes sense
  -> adjust the prompt or code
  -> repeat

A lot of the direction came from human feedback first. We bounced ideas between ourselves, then used AI to corroborate, stress-test, or quickly prototype them.

We also did lightweight UAT with people around us, including Rohan's parents. That was useful because they reacted less to the technical novelty and more to whether the advice felt practical. Some features, like being able to mention timeline or salary expectations and have the app respond realistically, came from those conversations.

PyCon SG Influence

Several project choices were either inspired by or validated by sessions during PyCon SG, but mostly in a practical way: they gave language to problems we were already running into while building.

For example, the AI-agent and reliability sessions matched what we were already feeling: agents are useful, but only if they are bounded. That pushed us toward allowlisted tools, deterministic scoring, fallback behavior, and evidence payloads instead of a freeform career chatbot. Talks like "Merlions, Agents & Copilot: Trustworthy Python on Azure", "This Talk Was Generated by AI. Please Don't Trust It", and Anthony Tung's keynote on using tools without being used by them helped us frame that choice more clearly.

The data/API side also became more real once we worked with MyCareersFuture and Apify. External data was useful, but schema changes, blocking, caching, and source preservation mattered more than expected. That made sessions like "Designing Python APIs for Data You Don't Control" feel directly relevant rather than abstract.

The testing and evaluation workshops reinforced the same point from another angle. If we were going to claim that the data decides, we needed tests and failure cases around scoring, parser guardrails, and fallback behavior. "Efficient Python Testing: Leveraging Parametrization, Fixtures, Monkeypatching and Marks" and "Do you know how well your model is doing? Evaluate your LLMs" were useful references for that.

The agent-process sessions also validated something more mundane but important: shared instructions matter when multiple humans and agents are touching the same repo. "SKILL.md is the SOP your AI agent never had" lined up with our use of AGENTS.md, DECISIONS.md, and collaboration logs. Those docs were not process theatre; they kept the project moving.

Finally, the Python tooling discussions lined up with how we treated Python in this project: not flashy, just reliable. "Adopting uv and pyproject.toml for mono-repo: Challenges and Approach" matched our choice to keep setup reproducible with uv and pyproject.toml.

Google Cloud was also a practical fit for hosting and scheduled refreshes. We used Cloud Run and Cloud Scheduler because they were straightforward for this kind of small app plus recurring data job.

Local Development

uv sync
cp .env.example .env
# add OPENAI_API_KEY in .env if using live AI features

uv run python scripts/01_ingest_real_data.py
uv run python scripts/02_build_skill_embeddings.py

# Optional: update the local MyCareersFuture SQLite history cache.
uv run python scripts/05_update_live_jobs.py --db var/mcf_jobs.sqlite --log logs/mcf_new_jobs.jsonl --limit 100 --max-pages 10

# Optional: import the latest open MyCareersFuture SQLite snapshots into DuckDB.
MCF_SQLITE_PATH="var/mcf_jobs.sqlite" uv run python scripts/03_ingest_jobs_snapshot.py

# Optional after MCF import: semantically map MCF skill strings to SkillsFuture Unique Skills.
uv run python scripts/04_build_mcf_skill_crosswalk.py

# Fold the embeddings into the DuckDB so it's the single, self-contained runtime artifact.
uv run python scripts/06_embeddings_to_duckdb.py

# Map the full course catalogue to framework skills semantically.
uv run python scripts/07_build_course_skill_crosswalk.py

npm --prefix frontend install
npm --prefix frontend run build

uv run uvicorn skillgraphsg.api.main:app --host 127.0.0.1 --port 8000

Open:

http://127.0.0.1:8000

For active UI work, use the Vite dev server instead:

uv run uvicorn skillgraphsg.api.main:app --port 8000
npm --prefix frontend run dev

The backend serves the built Vite app from frontend/dist/ at / and the API at /api. If the build is missing, it falls back to the legacy vanilla UI in static/.

Request-time OpenAI calls are opt-in for demo latency. Set SKILLGRAPH_REQUEST_OPENAI=1 for live profile parsing, scanned-PDF extraction, and request-time semantic matching. Also set SKILLGRAPH_LIVE_ORCHESTRATOR=1 to exercise the full multi-call tool loop. Set SKILLGRAPH_STRICT_OPENAI=1 for strict full-stack smoke testing with no silent fallback.

Optional LinkedIn enrichment uses Apify before profile parsing. Add a LinkedIn profile URL on onboarding, then configure:

APIFY_TOKEN="..."
APIFY_LINKEDIN_ACTOR_ID="get-leads/linkedin-scraper"
APIFY_LINKEDIN_INPUT_JSON='{"mode":"profiles","urls":["{url}"]}'

If Apify is not configured or the Actor fails, CareerSphere skips LinkedIn evidence and continues with the pasted/resume context.

MotherDuck Live Jobs

The local skillgraph.duckdb can be promoted to a MotherDuck database with scripts/07_upload_duckdb_to_motherduck.py. The uploader preserves the base SkillsFuture tables, embeddings, crosswalks, and current live jobs in one database.

For scheduled refreshes, build and deploy the Cloud Run job:

export GOOGLE_CLOUD_PROJECT="..."
export MOTHERDUCK_DATABASE="skillgraph"
export MOTHERDUCK_SECRET="motherduck-token"
scripts/deploy_mcf_ingest_cloud_run_job.sh
gcloud run jobs execute mcf-live-jobs-ingest --region asia-southeast1 --wait

The Cloud Run job runs scripts/08_update_live_jobs_motherduck.py, fetches recent MyCareersFuture pages, and writes directly to MotherDuck jobs_snapshot and job_skills. MCF skills that do not have a confident crosswalk remain market-only with unique_skill_id = NULL.

The deploy script also creates/updates a Cloud Scheduler trigger. In GCP, the current job is mcf-live-jobs-ingest, scheduled as mcf-live-jobs-ingest-daily at 0 2 * * * (Asia/Singapore) against md:skillgraph.

When MOTHERDUCK_TOKEN is set, normal app/runtime calls to skillgraphsg.engine.real_data.connect() read from md:${MOTHERDUCK_DATABASE} by default. Set SKILLGRAPH_USE_MOTHERDUCK=0 for local-only development or tests.

Tests

uv run pytest

The tests cover scoring behavior, deterministic tool boundaries, fallback behavior, semantic matching fallback, API shape, MCF skill crosswalk behavior, portfolio crawling, scanned-PDF extraction guardrails, job ranking, and evidence requirements.

Strict live smoke test:

uv run python scripts/smoke_strict_filtered_jobs.py

Known Limitations

  • Some job data is cached for demo stability.
  • Scraping and enrichment can fail or get blocked.
  • Course matching is useful but still imperfect.
  • The live course API / course-detail path is supposed to be maintained, but some returned courses still lead to 404s.
  • LLM parsing is conservative and intentionally capped at Level 4.
  • The 3D sphere helps exploration, but still needs the surrounding explanations to be useful.
  • The repo is a hackathon repo and shows the speed of iteration pretty honestly.

Credits

Built by Rohan and Kieran.

Thanks to the PyCon SG organizers, speakers, sponsors, partners, hackathon community, and everyone who gave feedback in person or through Telegram.

Special thanks to Rohan's family as well. While we were talking through the hackathon at Rohan's house, they pointed out that a timeline would be a good way to wrap the recommendation into something practical, which helped shape how we thought about next steps.

About

CareerSphere — grounded career exploration on Singapore's official SkillsFuture and live jobs data (PyCon SG 2026 hackathon).

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages