Skip to content

Annotation workbooks: name column doesn't match the English source; "change" replaced by a street address #14

Description

@bertfil

The name column in data/seatau/annotations/annotation_*.xlsx doesn't match the English source it's supposed to hold. Entities look like they've been swapped per-locale, and one of the swaps has hit a common English word.

tasks.json::0/user_scenario/instructions/known_info:

data/tau2/domains/retail/tasks.json : You are Yusuf Rossi in zip code 19122.
annotation_id.xlsx                  : You are Gatot Situmorang in zip code 15182.
annotation_th.xlsx                  : You are ณัฐจศักดิ์ เช้าวันดี in zip code 66360.

The bad one: the verb "change" gets replaced with a street address — Gg. Tubagus Ismail No. 385, Sorong, Sulawesi Tengah 34824 in id, locale-equivalents in the others. 2,558 rows in annotation_id.xlsx, mostly telecom_tasks:

"this action can only be called once, and will Gg. Tubagus Ismail No. 385,
 Sorong, Sulawesi Tengah 34824 the order status to 'pending (items modified)'"

"...the agent asks for confirmation, suddenly Gg. Tubagus Ismail No. 385,
 Sorong, Sulawesi Tengah 34824 your mind and ask to only..."

annotation_tl.xlsx does the same to ID → Anita Ville.

Comparing every tasks.json:: row in retail_tasks against the committed source:

workbook divergent / 473
id 131 (27.7%)
th 113 (23.9%)
vi 130 (27.5%)
tl 141 (29.8%)
zh 113 (23.9%)

Same token corrupted in all five but a different address each time, so detection probably ran once over English and substitution per-language. Couldn't tell for sure — I didn't find the script that does this in the repo.

Simulation results look unaffected; the shipped tasks.json and db.json are clean.

repro
# pip install openpyxl requests
import io, json, openpyxl, requests

RAW = "https://raw.githubusercontent.com/SEACrowd/SEATauBench/main"
tasks = requests.get(f"{RAW}/data/tau2/domains/retail/tasks.json").json()
wb = openpyxl.load_workbook(io.BytesIO(
    requests.get(f"{RAW}/data/seatau/annotations/annotation_id.xlsx").content),
    read_only=True, data_only=True)

ws = wb["retail_tasks"]
rows = list(ws.iter_rows(values_only=True))
hdr = [str(c) for c in rows[0]]
rec = dict(zip(hdr, rows[3]))

print("source  :", tasks[0]["user_scenario"]["instructions"]["known_info"])
print("workbook:", rec["name"])

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions