This repository turns real AppleSupport conversations from the Customer Support on Twitter dataset into a cautious support-agent prototype. It classifies one customer tweet into an operational intent, retrieves a historically similar AppleSupport response as a draft, and decides whether the case can be auto-handled or needs a person.
The project treats proof as more important than automation. It currently has no valid headline metric: the prior local gold/evaluation pass was model-assisted and is excluded from submission evidence. See REVIEWER_INTEGRITY.md.
The input is an incoming customer tweet. The agent returns:
- One of nine intents:
account_security,billing_payments,device_hardware,software_setup,connectivity,icloud_data,orders_delivery,app_store_media, orother. - A draft reply retrieved from historically similar AppleSupport cases, with evidence IDs for audit.
auto_handleorescalate, with a stated reason. Fraud, account access, payment disputes, privacy/data loss, legal language, lost/stolen devices, high-frustration complaints, and ambiguous messages escalate.
It does not access accounts, authenticate customers, change orders, issue refunds, make policy promises, or send a reply to Twitter.
The source file is local: data/twcs.csv (or set TWCS_PATH). npm run prepare makes two streamed passes over it:
- Collect AppleSupport tweets that directly reference an earlier tweet through
in_response_to_tweet_id. - Find those inbound customer tweets and retain non-empty customer/reply pairs.
The output is artifacts/apple_pairs.csv. The extraction uses direct reply edges rather than reconstructing entire Twitter threads, because the raw source contains missing and multi-ID links. Each extracted row receives a provisional weak label used only for the initial model and stratified sampling; it is not a gold label.
- Node.js 18 or later. No dependency installation is required.
data/sample_twcs.csvis included for a real-data extraction smoke test. The full corpus is available from the Kaggle dataset; seeDATA_PROVENANCE.mdbefore downloading or redistributing it.- An OpenAI API key is optional and is only required for
npm run judge.
| Command | Purpose |
|---|---|
npm run prepare |
Stream-extract direct AppleSupport/customer pairs from the raw corpus. |
npm run prepare:sample |
Extract the committed 180-pair real-data smoke sample to artifacts/sample_apple_pairs.csv. |
npm run make-sample |
Regenerate the committed smoke sample from a locally extracted full corpus. |
npm run make-gold |
Create a deterministic 200-row candidate gold sheet. |
npm run assist-gold |
Create separate model suggestions and uncalibrated confidence scores. Does not write gold labels. |
npm run review-gold -- ... |
Run an interactive human-labeling session and save a reviewer-specific CSV. |
npm run make-overlap -- ... |
Create a deterministic, empty 40-case second-review sample. |
npm run compare-reviewers -- ... |
Create a disagreement report from two approved reviewer files. |
npm run finalize-gold -- ... |
Copy explicitly human-approved rows into the canonical artifacts/golden_eval.csv. |
npm run validate-gold |
Verify count, schema, source linkage, uniqueness, and approved labels. |
npm run evaluate |
Generate measured metrics, predictions, and confusion matrices. Requires a valid final gold file. |
npm run update-report |
Insert only generated evaluation metrics into report/REPORT.md. |
npm run judge |
Run the optional LLM reply-quality rubric. Requires OPENAI_API_KEY. |
npm run judge-agreement |
Calculate judge-human agreement from independently entered human scores. |
npm run demo |
Run a sample agent response. |
npm test |
Run parser, routing, and model smoke tests. |
The committed smoke path runs without the full corpus and demonstrates extraction on 180 real AppleSupport pairs. It does not reproduce a headline metric. A valid headline metric remains blocked on independent human labels and judge-human agreement.
# From C:\Users\Admin\OneDrive\Desktop\hiver-support-agent
npm test
npm run prepare:sample
npm run demo:sampleTo run the full local pipeline, download the source through Kaggle under its stated license and place it at data/twcs.csv:
$env:TWCS_PATH = 'C:\path\to\twcs.csv'
npm run prepareThe original assignment’s “reproduce headline results in under 15 minutes” requirement is not currently satisfied because no valid headline result exists. Do not claim that requirement is met until the reviewer-integrity steps below are complete.
The key artifacts after this stage are:
artifacts/apple_pairs.csv— local training/retrieval pairs.artifacts/golden_eval_draft.csv— 200 candidate examples with empty human fields.artifacts/golden_eval_reviewer_assistance.csv— separate suggestions, confidence bands, and review priorities. Itsgold_*columns remain blank.
Read evaluation/LABELING_GUIDE.md and REVIEWER_INTEGRITY.md before annotation. Labeler A must start from the blinded candidate sheet, not the model-assistance sheet.
Start a personal review file. The command never overwrites an existing file unless --resume is specified:
npm run review-gold -- --labeler "Reviewer A" --output artifacts/reviews/reviewer_a.csvDo not use model suggestions or the fast-review bulk-approval path to establish the evaluation gold set. They are appropriate for triage only and invalidate independent-label claims.
Create that overlap and compare the two reviewers without copying either person's labels into the other file:
npm run make-overlap
npm run review-gold -- --input artifacts/reviews/reviewer_b_overlap_draft.csv --labeler "Reviewer B" --output artifacts/reviews/reviewer_b.csv
npm run compare-reviewers -- --first artifacts/reviews/reviewer_a.csv --second artifacts/reviews/reviewer_b.csvAn adjudicator resolves disagreements recorded in artifacts/reviews/label_agreement.csv. To revisit Reviewer A's file into a separate adjudicated copy, add --revisit-approved to review-gold; the script still requires a human to enter and approve every final decision.
After a human has explicitly entered labels and marked 150–250 rows approved, create the canonical file:
npm run finalize-gold -- --input artifacts/reviews/reviewer_a.csv
npm run validate-goldvalidate-gold checks that every final example still matches apple_pairs.csv, has a unique ID, belongs to the defined taxonomy, includes a route reason and labeler, and has the required approved count. It does not create labels.
Use the faster reviewer when a person is comfortable explicitly accepting or correcting the displayed model suggestion. It reads only artifacts/golden_eval_reviewer_assistance.csv, creates artifacts/reviews/ if needed, and resumes from the first unapproved row if the output file already exists.
npm run review-gold-fast -- --labeler "Reviewer A" --output artifacts/reviews/reviewer_a_fast.csvFor every case it displays the customer message, suggested intent/route/reason, numeric confidence, confidence band, and priority. Enter one of:
yto explicitly approve the displayed suggestion.nto enter the intent, route, and route reason yourself; the edited decision is then approved under your reviewer name.sto leave the row unapproved and move on.qto save immediately and stop.yato request bulk approval of remaining cases whose confidence band ishighand numeric confidence is at least0.95. The command shows the eligible count and requires a secondyconfirmation before changing rows.
Every action saves the reviewer output. The fast command does not create artifacts/golden_eval.csv or auto-run evaluation. Its output must not be used as independent gold evidence because it exposes and can bulk-approve model suggestions.
Only after a newly created, independently human-labelled gold set passes npm run validate-gold:
npm run evaluate
npm run update-reportEvaluation excludes every canonical gold ID from both training and retrieval. It writes:
artifacts/evaluation_results.json— metrics, per-class results, and all confusion matrices.artifacts/evaluation_summary.md— concise summary table.artifacts/intent_confusion_matrix.csvandartifacts/routing_confusion_matrix.csv— reviewer-readable matrices.artifacts/predictions.csv— individual predictions and retrieved evidence IDs.
The evaluation compares the agent with a majority-intent baseline and a keyword-rule baseline. Routing is policy-based, not a separately supervised model; this is disclosed in the output notes and report.
The optional LLM judge uses a fixed temperature-zero rubric for relevance, groundedness, safety, helpfulness, tone, and overall quality:
$env:OPENAI_API_KEY = '...'
$env:JUDGE_LIMIT = 200
npm run judgeTwo support-literate humans must independently score at least 30 replies on the same 1–5 rubric. Enter adjudicated scores in human_overall in artifacts/llm_judgments.csv, then run:
npm run judge-agreementThe resulting artifacts/judge_agreement.json reports exact agreement and Cohen’s kappa only when there are at least 30 paired human/LLM scores. Read evaluation/HUMAN_JUDGE_PROTOCOL.md for the protocol. Until that artifact is available, judge-human agreement is not established.
report/REPORT.md— problem framing, results section, limitations, failure hypotheses, misleading-headline discussion, and next steps.DECISION_LOG.md— non-obvious design decisions and rationale.SUBMISSION_CHECKLIST.md— requirement-by-requirement current status.FINAL_STATUS.md— completion state and exact next commands.
Raw tweet data and all evaluation artifacts are ignored by Git by default. Commit a canonical gold set or result only after it has passed the integrity process in REVIEWER_INTEGRITY.md, and check the original CC BY-NC-SA 4.0 license in DATA_PROVENANCE.md. Keep raw corpus data, extracted pairs, predictions, reviewer workbooks, and detailed LLM judgments local unless redistribution is permitted. Do not treat retrieved historical replies as current Apple policy.