Skip to content

Latest commit

 

History

100 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SkillEye: Skeleton-Based Motion Quality Analysis for Racket Sports

Submission for the 8th CTCI Science and Technology Creativity Competition 2026 (Sports theme).

Abstract

Existing racket-sport analytics apps (e.g. SwingVision) center on ball tracking and match outcome — where the ball landed, who won the point. SkillEye instead analyzes the player's movement, extracting a normalized 2D skeleton from swing video with RTMPose and modeling it with a Spatio-Temporal Graph Convolutional Network (ST-GCN), toward the goal of a per-swing motion-quality score with joint/phase-level error detection and rule-based correction feedback. This addresses both the "training methods" and "injury prevention" sub-themes of the competition through one shared representation.

This document reports the technical work completed to date on the THETIS tennis dataset: a validated pose-extraction pipeline (1,980/1,980 clips, 0 failures), a 6-class stroke-type classifier (81.7% ± 4.9%, 5-fold subject-disjoint cross-validation), and a beginner-vs-expert motion-quality proxy classifier (82.4% ± 3.8%). Both are real trained results on held-out subjects, not projected targets. A rule-based v1 of that quality-score system — phase detection, per-joint deviation from expert-clip templates, and generated correction suggestions — is implemented and demonstrated through an accompanying Streamlit UI. Two later additions (Sections 2.8–2.10) build on this: a correlated version of the scorer that checks each joint against the others rather than in isolation, two rules that check for specific mistakes published biomechanics research links to lower skill, and a racket-mounted sensor that has moved from a finalized PCB design to a built, tested board streaming real motion data. A third addition (Section 2.11) closes the loop those first two open: the full pipeline -- capture, pose extraction, both scorers, the skill-level rules -- run end to end on footage this team recorded itself, not just THETIS, with real, specific findings (Section 3.5). Section 4 discusses what these results do and do not establish, and Section 5 lays out the remaining work toward the full quality-score system.

1. Introduction

1.1 Motivation

Amateur and intermediate racket-sport players typically improve technique through in-person coaching, which is expensive, geographically limited, and infrequent relative to how often players practice alone. Video-based swing analysis could close this gap, but existing consumer tools (SwingVision and similar) are built around ball-tracking and scorekeeping — they answer "did the shot land in" rather than "what did the player's body do." Movement quality is the input a coach actually reasons about (elbow position at contact, hip rotation timing, follow-through), and it is not derivable from ball trajectory alone.

1.2 Related Work

The technical approach follows two established lines of work rather than proposing new architectures: RTMPose for real-time 2D human pose estimation, and ST-GCN (Yan, Xiong & Lin, 2018) for skeleton-based action recognition, which models a keypoint sequence as a spatio-temporal graph — spatial edges follow the human skeleton, temporal edges connect the same joint across frames. The novelty in SkillEye is the application: using this representation not just to classify what stroke was performed (a mature task), but as the foundation for scoring how well it was performed, which is comparatively unexplored for racket sports at consumer scale. The THETIS dataset itself (Gourgari et al., CVPR 2013 workshop) was built for action recognition, not quality assessment — Section 2.1 and 4.1 discuss what that means for how it can and cannot be used here.

A useful point of comparison is Wagner's tennis serve dataset (2024, jasnwag/tennis_serve_dataset), which takes a near-opposite set of trade-offs to THETIS: broadcast footage of the 2024 US Open, 109 professional players, 5,966 serves, with 3D pose (RTMPose → MotionBERT lifting → DTW-aligned) rather than 2D, and — notably — an actual serve-quality classification task (84.0% reported), where THETIS offers only the subject-level beginner/expert proxy used in Section 2.4/3.2. Table 2 summarizes the contrast:

Table 2. Dataset comparison.

THETIS (this work) Wagner tennis serve dataset
Source controlled gym recording (Kinect) broadcast video, 2024 US Open
Camera angle near-frontal broadcast (side/elevated)
Pose dimensionality 2D (COCO-17) 3D (COCO-17, MotionBERT-lifted)
Subjects 55 (31 beginner, 24 expert) 109 professional players
Stroke coverage 12 categories, all strokes serves only
Clips 1,980 5,966 serves
Quality ground truth none (subject-level skill label only) serve-quality labels available
License dataset's own terms CC BY 4.0

The two datasets are complementary rather than substitutable for this project: THETIS's frontal angle and 2D pose match what a consumer phone camera can realistically capture (the deployment target), while the serve dataset's 3D pose and side/elevated broadcast angle are closer to what the biomechanical joint-angle rules in Section 4.7 of the proposal actually need, and its existing quality labels are exactly the kind of ground truth Path A/B (Section 5) is trying to establish independently. It is treated here as a comparative reference and a candidate source for cross-dataset validation once the 2D pipeline has a 3D-compatible variant, not (yet) as training or evaluation data for the models in Section 3 — see Section 5 for that as a concrete next step.

1.3 Contributions

  1. A validated, resumable pose-extraction pipeline from raw racket-sport video to normalized, tracked, single-subject skeleton sequences (Section 2.2).
  2. A stroke-type classifier trained from scratch on the extracted skeletons, evaluated with 5-fold subject-disjoint cross-validation rather than a single split (Section 3.1).
  3. Evidence, from a dedicated experiment, that motion quality (not just stroke type) is learnable from this same skeleton representation — using THETIS's beginner/expert metadata as a weak proxy label ahead of real coach-rating ground truth (Section 3.2).
  4. A working, end-to-end swing-quality feedback loop, not just a classifier: phase detection, per-joint deviation from expert templates, rule-based correction suggestions, and an optional LLM layer that rewrites those same flagged deviations into a coaching paragraph — demonstrated through an accompanying Streamlit UI (Section 2.6).
  5. A cross-disciplinary hardware + software effort toward a wearable sensing device: a finalized (rev2.1), PCB-laid-out IMU sensor board designed in parallel with the modeling work in this document, plus a software fusion architecture built to consume its data the moment real recordings exist (Section 2.7).
  6. An honest accounting of what this dataset cannot yet support (Section 4.1) and a concrete plan for closing that gap (Section 5).

2. Methodology

2.0 System Overview

Before the section-by-section detail below, this is how the pieces fit together end to end — from a raw swing video to what a user actually sees. Solid blue boxes are built and validated with real trained results (Sections 2.1–2.6, 3); the dashed amber subgraph is the sensor-fusion prototype (Section 2.7), which is real code and a real hardware design, but not yet trained on real sensor data.

System architecture: swing video through RTMPose, both ST-GCN classifiers, quality scoring, optional LLM explainer, the Streamlit UI, and the sensor-fusion prototype branch

2.1 Dataset

THETIS (Gourgari et al., 2013) provides 1,980 RGB video clips across 12 tennis action categories, performed by 55 subjects (p1–p31 labeled beginner, p32–p55 labeled expert per the dataset's own metadata), recorded at ~17–19 fps, 640×480, with a Kinect camera facing the subject (near-frontal, not side-on). Only the VIDEO_RGB subset is used — VIDEO_Skelet2D/3D uses Kinect's joint definition (not COCO-17, the format the deployed pipeline actually produces) and covers only 1,217 of 8,374 sequences, so it would not be representative even if remapped.

Class composition after the merge described in 2.3 (Table 1):

stroke THETIS sub-categories merged clips
backhand backhand, backhand2hands, backhand_slice 495
forehand forehand_flat, forehand_openstands, forehand_slice 495
backhand_volley backhand_volley 165
forehand_volley forehand_volley 165
serve flat_service, kick_service, slice_service 495
smash smash 165
total 1,980

By skill label: 1,116 clips from 31 beginner subjects, 864 clips from 24 expert subjects.

Measured limitations relevant to this project (checked directly on sample frames, not assumed from the dataset paper): the ~17-19fps frame rate is below what the proposal's biomechanical rules assume is needed to resolve fast strokes (e.g. serves), and the near-frontal camera angle does not capture the side-on view that joint-angle-based error-detection rules require. THETIS is therefore treated here as suitable for pretraining the motion representation and for stroke/skill classification (both established, well-precedented tasks — this is what Sections 2.3–2.4 and 3 use it for), but not as a substitute for purpose-collected side-view footage when the actual per-swing quality-score model is built (Section 5).

2.2 Pose Extraction Pipeline

Each clip is processed frame-by-frame with RTMPose (rtmpose-s_simcc-body7, lightweight mode, COCO-17 keypoints) for pose estimation and YOLOX-tiny for person detection, both run via rtmlib/ONNX Runtime. Frames may contain more than one person (other players/staff visible in the gym), so a dedicated tracking step selects a single, temporally consistent subject:

  1. Primary-subject selection — on the first frame, the person with the largest visible bounding box (among detections with ≥3 confidently-visible joints) is chosen as the subject. On subsequent frames, the detection with the nearest hip-center to the previous frame's selection is kept, which prevents identity swaps between the subject and any other person entering the frame (largest-bbox alone is insufficient once multiple people are present).
  2. Confidence-based interpolation — per joint, frames below a 0.3 confidence threshold are linearly interpolated from the nearest confident frames on either side (no extrapolation at clip boundaries). Clips where the primary subject is confidently visible in under 50% of frames are dropped rather than interpolated across too large a gap.
  3. Normalization — each frame is re-centered on the hip midpoint and scaled by shoulder-to-hip (torso) length, so the model learns the shape of the motion rather than the subject's distance from or position relative to the camera.

Result: 1,980/1,980 clips extracted successfully, 0 dropped, 0 failed, ~1.4s/clip. Code: ml/skilleye/skeleton_pipeline.py (steps above), ml/skilleye/batch_extract.py (CLI batch runner, resumable — skips clips that already have output on disk, so an interrupted run can restart without redoing completed work).

2.3 Stroke Classification Model

THETIS's 12 action categories are merged into 6 target classes (Table 1) reflecting the strokes the proposal is actually built around — serve merges the three service sub-types (genuinely one stroke with different spin), but forehand_volley and backhand_volley are kept separate rather than merged into a single "volley" class, since they are visually and kinematically distinct motions (an earlier iteration that merged them made that class disproportionately hard to classify — see Section 4.1).

Architecture: a compact ST-GCN — 6 spatio-temporal blocks operating on the COCO-17 skeleton graph (symmetric-normalized adjacency, self-loops included), channel widths 32→64→128→256, global average pooling, linear classification head. Input is a 4-channel tensor per joint per frame: 2D position plus its temporal finite-difference (velocity), giving the model direct access to motion speed/direction rather than only inferring it across pooled layers. Each clip is time-resampled to a fixed 64 frames (linear interpolation) so variable clip lengths (THETIS clips range from roughly 2 to 8+ seconds) can be batched. Code: ml/skilleye/stgcn_model.py.

Training: Adam (lr 1e-3, weight decay 1e-4), cosine-annealed learning rate, class-weighted cross-entropy (classes are imbalanced — see Table 1), 80–100 epochs, batch size 32. With only 55 subjects total, label-preserving data augmentation is applied at train time: left-right mirroring (negate the x-coordinate and swap left/right joint indices — a mirrored forehand is still a forehand, just as a left-handed player's would look), random temporal cropping (retain a random 85–100%-length contiguous sub-window before resampling), and small Gaussian coordinate jitter (σ=0.02, in normalized units). Code: ml/skilleye/train_stroke_classifier.py.

2.4 Beginner-vs-Expert Motion-Quality Proxy Model

To test whether motion quality — not just stroke identity — is learnable from this representation, a second model is trained on THETIS's beginner/expert subject-level label (Section 2.1) as a weak proxy for swing quality, ahead of collecting real coach ratings (Section 5). Two approaches were tried, reported both for transparency:

  • v1 — hand-crafted features + logistic regression (ml/skilleye/beginner_expert_check.py): 9 interpretable per-clip features (mean/std swing speed, jerk, elbow/knee joint-angle variance, wrist extension range) fed into a linear classifier.
  • v2 — ST-GCN (ml/skilleye/train_beginner_expert_stgcn.py): identical architecture and augmentation to Section 2.3, but with a binary beginner/expert output head, trained directly on the raw skeleton sequence rather than hand-summarized statistics.

2.5 Evaluation Protocol

All splits are subject-disjoint: no subject's clips appear in both the training and validation sets. THETIS clips from the same subject are highly correlated (same body, same camera, same session), so a clip-level random split would leak subject identity into validation and overstate accuracy — the model would partly be recognizing people, not strokes or skill.

Headline numbers use 5-fold subject-disjoint cross-validation (ml/skilleye/cross_validate.py): subjects are split into 5 folds, stratified so each fold's held-out group keeps a beginner/expert mix; each fold trains a fresh model on the other 4 folds' subjects (~44) and evaluates on the held-out fold's subjects (~11). This is reported as mean ± standard deviation across folds, rather than a single split's number, because with only 55 subjects a single split's held-out accuracy depends materially on which subjects happened to be held out — the fold-to-fold spread reported in Section 3 is itself informative about that variance.

2.6 Quality Scoring System

The proposal's central promised feature -- a per-swing motion-quality score with per-joint/phase error detection and rule-based correction suggestions (Section 4.7) -- is implemented here as a rule-based comparison against expert-clip statistics, not a learned regression model. No coach-rating ground truth exists yet (Path A/B, Section 5), so a supervised quality-score model cannot be trained honestly; this is an explicitly-scoped interim v1, in the same spirit as the v1->v2->v3 iterations already described for the classifiers above.

Example poses at the detected contact frame, one per stroke class

Phase detection (ml/skilleye/quality/phases.py): each clip's dominant wrist (whichever moves more overall -- a proxy for the hitting arm, since THETIS doesn't label handedness) is used to find the contact frame (peak wrist speed, a standard swing-analysis heuristic), splitting the clip into three phases: backswing, a short window around contact, and follow-through.

Joint angles (ml/skilleye/quality/angles.py): per phase, six flexion angles (left/right shoulder, left/right elbow, left/right knee) plus a trunk-rotation proxy (shoulder-line vs. hip-line angle) are averaged into one scalar per (phase, joint) — 7 joints × 3 phases = 21 values per clip. Shoulder angle (angle at the shoulder between hip and elbow) was added in a later pass (Section 2.8) so Elliott (2006)'s "shoulder" kinetic-chain contributor has a tracked joint to correspond to; it is a 2D flexion/abduction proxy, not a reconstruction of true 3D shoulder rotation, which a single camera cannot measure.

Expert templates (ml/skilleye/build_expert_templates.py, run once offline): for each stroke class, the mean and standard deviation of each (phase, joint) scalar across that class's expert-labeled clips, computed only from the training side of this project's standard subject-disjoint split (Section 2.5) -- never from the held-out validation subjects, which remain available as genuinely unseen demo inputs.

Scoring (ml/skilleye/quality/score.py): a query clip's (phase, joint) scalars are z-scored against its predicted stroke's template; |z| > 1.5 is flagged and mapped to a fixed coaching-tip sentence, and the overall 0-100 score is a monotonic function of mean absolute deviation. SCORE_SCALE/FLAG_THRESHOLD are sanity-checked (Section 3.3), not formally calibrated -- that calibration is exactly what Path A/B ground truth would enable.

Demo UI (ml/skilleye/app.py, Streamlit): pick a stroke category and a sample clip (restricted to the held-out validation subjects, so every demo score is computed against a template that never saw that subject) and see the skeleton, the existing stroke classifier's prediction, the quality score, the per-phase/joint table, and the correction suggestions together. Run with streamlit run app.py from ml/skilleye/.

AI-generated explanation (optional): a "Generate AI explanation" button in the demo UI rewrites the already-flagged deviations above into one natural coaching paragraph via an LLM (NVIDIA's build.nvidia.com API). The LLM is only ever given the (phase, joint, z-score) rows the rule-based system already flagged — it never inspects raw angles or introduces a new diagnosis — so this is a communication layer over the existing, smoke-checked scoring system, not a new source of truth. It requires an NVIDIA_API_KEY environment variable; without one (or if the API call fails for any reason), the button falls back to a short warning and the rule-based suggestions list above remains the result, so the demo never depends on network access to function. Design: docs/superpowers/specs/2026-07-23-llm-correction-explainer-design.md.

2.7 Sensor-Fusion Extension (Prototype)

A team discussion identified a concrete limitation already described in Section 4.1: a single frontal 2D camera cannot see racket-face angle or wrist-snap dynamics at contact — exactly the kind of high-frequency, off-camera-plane signal a racket-mounted IMU (accelerometer + gyroscope) would capture directly. This section covers both halves of that work: the hardware (designed by a teammate, in parallel with everything else in this document) and the software fusion prototype that will eventually consume its data.

Hardware — rev2.1, finalized (designed by @kyriosaa): the design went through two iterations. Rev1.0/1.1 was a fully custom board — an ESP32-C6-WROOM-1 microcontroller, an ST LSM6DSO 6-axis IMU, and discrete TP4056/DW01A/FS8205A charging-protection ICs, refined once (rev1.1: added a VBUS sense divider). Rev2.0/2.1 then deliberately traded that custom-everything approach for off-the-shelf breakout modules — a GY-521 (the standard MPU6050 accelerometer+gyroscope breakout — the exact sensor originally proposed in the team's first discussion of this feature, Section 2.7's origin) on a Seeed Studio XIAO microcontroller, with a TP4056 charging breakout — lower assembly risk for a first real board, at the cost of a larger footprint than the fully custom rev1.0. Full PCB layout is done (skilleye_2.0.kicad_pcb), and the project now includes a PCBWay logo asset, indicating the design is being prepared for fabrication. Full KiCad project: hardware/skilleye_2.0/; rendered schematic below, source at docs/schematics/2.0/rev2.1/. The earlier rev1.0/1.1 custom-board files remain in hardware/skilleye_1.0/ / docs/schematics/1.0/ for reference.

Skilleye 2.0 schematic (rev2.1, finalized)

Where this board would sit on the racket, and what its two sensing modalities are meant to pick up (both are concept diagrams — no physical unit has been mounted yet):

Placement

Sensor placement concept

Sensing mechanism

Sensor mechanism concept

Update — the board is now built and streaming real data. What follows was true when the design above was finalized, but is no longer the current state: firmware has been written (hardware/firmware/firmware.ino — the XIAO ESP32-C6 + GY-521 hosts its own WiFi network and streams accelerometer/gyroscope readings at 200 Hz the moment a laptop connects), a computer-side client exists (hardware/client/imu_client.py), and real recordings exist for backhand/forehand/serve (with and without ball contact) plus one volley take. A live-monitoring dashboard (hardware/client/live_dashboard.py, Streamlit) was built and tested against the physical device on real WiFi, confirming the full sensor → firmware → recording pipeline works end to end — see Section 2.10. The fusion model in this section still trains on synthetic data only (below) — swapping in real recordings for that specific model is the one piece of this section still pending, not the whole hardware effort.

Software architecture (ml/skilleye/imu_fusion.py): STGCN is refactored to expose its pooled pre-classifier features via extract_features() (ml/skilleye/stgcn_model.py, backward compatible — every existing caller is unaffected). A small IMUEncoder (1D-CNN over a 6-channel accelerometer+gyroscope stream) is fused with that skeleton branch via concatenation before one classification head (FusedBeginnerExpertModel), targeting the beginner/expert distinction specifically, since that is the axis this sensor is meant to inform.

Synthetic data, stated plainly: because no real sensor data exists yet, synthetic_imu_from_skeleton() derives a placeholder signal from the skeleton itself (wrist acceleration, forearm angular velocity) rather than from a real sensor. This signal is, by construction, redundant with information the skeleton branch already has access to — so train_beginner_expert_fusion_prototype.py's output is a demonstration that the fusion architecture and training loop work end to end, not an accuracy result, and its numbers must not be cited alongside Section 3.2's cross-validated 82.4% ± 3.8%.

Path to real data: a documented collection protocol (~100-200 Hz logging, a tap-based manual sync event between the video and IMU streams, resampled to the same fixed frame count skeletons already use) is in docs/superpowers/specs/2026-07-23-imu-fusion-prototype-design.md — written before either hardware revision existed, describing a generic MPU6050+ESP32 pairing; the rev2.x board above (GY-521/MPU6050 on a Seeed XIAO) matches that description closely, and the collection method (sampling rate, sync approach) applies unchanged either way. Swapping synthetic_imu_from_skeleton()'s call site for a real-data loader is the only code change needed once the board above is built and recordings exist.

2.8 Correlated Quality Scoring (Module A)

The next three subsections describe work added after the rest of this document, and are written in plainer language on purpose so they're easy to follow even without a machine learning background.

The scoring system described in Section 2.6 checks every joint on its own. It looks at one joint's angle, compares it only to that joint's own usual value, and flags it if it's too different. This misses something real: sometimes a joint looks "off" only because another joint pulled it along. For example, if the trunk rotates a lot during a swing, the elbow angle naturally shifts too — that isn't a mistake, it's just how a real swing moves. The old system couldn't tell the difference between "this joint is genuinely wrong" and "this joint moved because the rest of the body moved."

Module A fixes this by scoring each joint together with the others, not alone. Instead of asking "is this elbow angle unusual by itself?", it asks "is this elbow angle unusual, given what the rest of the body is doing right now?" This uses a statistics tool called a covariance matrix — a table that captures how pairs of joints move together across real expert swings — plus a formula that predicts one joint's expected value from the others' current values. The gap between the real value and that prediction is the new score.

This covariance matrix is built separately for each stroke type (a serve and a volley move very differently, so they shouldn't share one set of rules), using the same expert-labeled clips already used to build the original templates — 1,584 training clips across the full 1,980-clip dataset. For strokes with fewer expert examples (volleys and smash, ~55-57 clips), the matrix is nudged toward a small, clearly-labeled "prior" — a rough guess, based on Elliott (2006)'s description of how the shoulder, elbow, and trunk work together during a swing, used only to keep the statistics stable when there isn't much data, not presented as a precise number from that paper.

The output format is unchanged: still one score per (phase, joint), the same 21 values as before, and the existing correction messages still work — only how each score is calculated is different. Code: ml/skilleye/quality/correlation.py, score_clip_correlated() in ml/skilleye/quality/score.py. Design: docs/superpowers/specs/2026-08-14-correlated-zscore-module-a-design.md. Validation results: Section 3.4.

2.9 Skill-Level-Specific Rules (Module B)

Module A still only compares a swing to what experts typically look like. It has no idea what beginners specifically tend to do wrong — it can only say "this is unusual," not "this is the mistake beginners commonly make." Module B closes that gap with a small number of hand-written rules, each based on one specific finding from a published biomechanics paper that directly compared skilled and less-skilled players.

Two rules were built, both for the backhand volley specifically (the one stroke a suitable comparison study exists for):

  • Shoulder-pelvis twist reversal. Katsumi et al. (2026) found that skilled players keep their shoulders twisted the same way relative to their hips throughout the swing, while less-skilled players' twist direction flips between the backswing and the moment of contact — as if the twist needed for a good hit was created too late. This project's existing angle-tracking code could only measure the size of this twist, not its direction, so two new "signed" angle functions were added (signed_shoulder_pelvis_twist_series, signed_pelvic_rotation_series in ml/skilleye/quality/angles.py) specifically so a direction reversal can be detected at all.
  • Excessive pelvic rotation. The same paper found less-skilled players rotate their hips further than skilled players, at both the backswing and contact. The rule flags a swing when its hip rotation passes the midpoint between the two groups' typical values.

A third rule — using Aydin & Aydemir (2026)'s finding that lower, more controlled racket acceleration corresponds to higher skill in a volley (elite players use a "controlled, abbreviated swing" rather than muscling the shot) — was implemented as a standalone check, check_volley_swing_effort(), but is not yet wired into the main pipeline: it needs a real racket-acceleration reading from an actual ball-contact volley swing, and no such recording exists yet (Section 2.10 explains what does exist). A fourth candidate rule, for forehand/backhand/serve, was investigated (a paper comparing high-performance and intermediate forehand drives) but turned out to need motion-capture and muscle-sensor (EMG) hardware this project doesn't have, so it was not attempted.

Every rule here produces its own separate flag, kept apart from Module A's 21 scores, since that's a structurally different kind of check (comparing to both skill groups, not just to experts). Code: ml/skilleye/quality/skill_rules.py. Design: docs/superpowers/specs/2026-08-14-skill-level-rules-module-b-design.md.

2.10 Live IMU Monitoring Dashboard

Separately from the modeling work above, a small tool was built to watch the racket sensor's data live, on a computer screen, while the racket is actually being swung — useful both as a quick "is the hardware actually working right now" check and as something to show during a demo. It reuses the existing recording client's networking code, runs the WiFi connection on a background thread, and redraws two live charts (accelerometer and gyroscope) plus simple stats (samples per second, dropped samples, peak force) a few times a second. It is intentionally kept separate from the main quality-scoring demo, so a problem here can never affect the validated, already-working pipeline. Code: hardware/client/live_buffer.py (the data-buffering logic, unit-tested), hardware/client/live_dashboard.py (the Streamlit page — run with streamlit run live_dashboard.py from hardware/client/).

This tool was tested against the real, physical racket sensor: connected over its WiFi network, confirmed streaming at the firmware's target 200 Hz, and used to record a genuine 58-second volley take (11,659 samples, zero dropped) — the project's first volley recording (previous recordings covered backhand/forehand/serve only). That specific take was swung without a ball, so its numbers are not yet comparable to Aydin & Aydemir (2026)'s ball-contact-based reference (Section 2.9) — but it did prove the complete chain, from physical sensor through to a rule-checking function, runs correctly end to end. A recording made with a real ball is the next step before that specific rule's output can be trusted.

Synchronized video + IMU recording (hardware/client/sync_recorder.py, needs pip install opencv-python): starts the laptop's webcam and the racket's IMU stream together from one command, instead of two manually-launched programs, for collecting the real (not synthetic) data Section 2.7's fusion model needs. The IMU's own clock and the webcam's frame clock still have no hardware link between them, so both are anchored to the recording computer's wall clock at the moment each starts (written to <name>_alignment.json) — this removes the human delay of starting two programs by hand, but the tap-sync convention above remains the fine-grained sync point, not replaced by it. Its frame-capture and recording logic is unit-tested with a fake camera and the same mock-device pattern used for imu_client.py (Section 2.10 above); actually recording a full session — multiple people, multiple swings — is tracked in Section 5 as separate, larger work.

A desktop GUI version also exists (hardware/client/sync_recorder_gui.py, additionally needs pillow): Start/Stop buttons plus a live webcam preview while recording (a framing check, not live pose tracking — this project's pose estimation runs offline on saved clips elsewhere), and a native "Save As" dialog on Stop, so the take's name is chosen after seeing it was worth keeping rather than before recording starts. Its file-renaming logic is unit-tested; the Tk window itself isn't (no display/camera to drive it against in this dev environment) — same manual-QA convention as live_dashboard.py.

2.11 Real-Recording Capture & Evaluation Pipeline

Sections 2.1–2.6 validate this project's pipeline entirely on THETIS. This section closes that loop once: every stage — capture, pose extraction, both quality scorers, and Module B — run end to end on footage this team recorded itself, with its own hardware, not a public dataset.

Capture. One take per stroke (backhand, forehand, backhand volley, forehand volley, serve, smash) was recorded with hardware/client/sync_recorder_gui.py (§2.10): the laptop's own webcam and the racket-mounted IMU sensor, started together and anchored to a shared wall-clock reference (<name>_alignment.json), with the tap-sync convention kept as the fine-grained backup. All six takes tracked the subject with RTMPose at 99–100% frame confidence — as high as anything in the THETIS pipeline.

Extraction (ml/skilleye/extract_real_recordings.py): the same rtmpose-s_simcc-body7 model and skeleton_pipeline.clean_clip() normalization used for THETIS (§2.2), pointed at these six arbitrary-filename .mp4 files instead of THETIS's batch folder convention. No pipeline code changed — proof that the extraction pipeline generalizes past the one dataset it was built and validated on.

Evaluation (ml/skilleye/evaluate_real_recordings.py): each extracted skeleton is scored with both score_clip() (independent, §2.6) and score_clip_correlated() (Module A, §2.8) against the existing THETIS-trained templates; the backhand-volley clip is additionally checked with evaluate_backhand_volley_skill_rules() (Module B, §2.9); and each take's real IMU peak acceleration is checked with check_volley_swing_effort() (§2.9) for both volley clips. The contact frame (peak wrist speed, same heuristic as quality/phases.py) is rendered as a skeleton figure alongside the real IMU trace for that take — Section 3.5 shows these.

Two limits, stated plainly rather than left implicit: (1) the templates being scored against were built from THETIS's Kinect footage — a different camera, room, and distance than this laptop webcam, so absolute scores here demonstrate the pipeline working, not a calibrated judgment of technique (§4.2 already makes this distinction for every other quality-score number in this project); (2) this is one clip per stroke from one person, not a dataset — a pipeline smoke test, not a statistically powered evaluation.

3. Results

3.1 Stroke Classification

81.7% ± 4.9% held-out accuracy (6-way; majority-class baseline 16.7%), fold range 72.5%–86.1%.

Accuracy across iterations

Stroke classifier confusion matrix

class precision recall f1 support
backhand 0.79 0.90 0.84 495
forehand 0.87 0.79 0.83 495
backhand_volley 0.66 0.63 0.64 165
forehand_volley 0.68 0.67 0.68 165
serve 0.93 0.88 0.90 495
smash 0.73 0.78 0.75 165

serve is the strongest class — mechanically the most distinct motion (toss + overhead strike). The volley classes are weakest, and the confusion matrix shows why: forehand_volley is confused primarily with forehand, backhand_volley primarily with backhand — i.e. a volley shares arm-swing kinematics with its groundstroke counterpart, and the near-frontal 2D camera does not capture the footwork/court-position difference (volleys are played closer to the net) that would otherwise disambiguate them. This is discussed further in Section 4.1.

Training curves

3.2 Beginner vs. Expert

82.4% ± 3.8% held-out accuracy (2-way; majority-class baseline 54.6%), fold range 78.5%–89.7%. The v1→v2 progression:

iteration method held-out accuracy
v1 9 hand-crafted features + logistic regression, single split 58.6% (baseline 54.6%)
v2 ST-GCN on raw skeleton sequence, single split 76.0%
v3 ST-GCN, 5-fold cross-validation 82.4% ± 3.8%

Beginner vs expert confusion matrix

v1 established that population-level differences between beginner and expert clips are real and highly significant (Welch's t-test on all 9 features, p<0.02, most p<1e-10, n=1,980) — notably, experts showed higher swing speed, jerk, and joint-angle variance than beginners, not lower, i.e. expert swings read as more dynamic/forceful rather than smoother in this dataset, which should inform how a smoothness-based quality heuristic is weighted. But v1's linear model on those 9 summary statistics only reached 58.6% held-out accuracy — barely above baseline, meaning the signal existed in aggregate but wasn't usable per-subject with simple features.

v2 replaced the hand-crafted features with the same ST-GCN architecture used for stroke classification, trained directly on the raw skeleton sequence, and reached 76.0% on a single split. v3 cross-validated that result and found it was, if anything, an unfavorable split — the cross-validated mean (82.4%) is higher than the single-split number, with folds ranging 78.5%–89.7% (Table above; per-class breakdown: beginner precision 0.85/recall 0.83, expert precision 0.79/recall 0.82, summed over all 5 folds' held-out predictions).

3.3 Quality Scoring Smoke Check

No formal ground truth exists yet to validate quality scores against (Section 2.6) -- this instead checks the one directional claim the system must satisfy to be credible: held-out expert clips should score higher on average than held-out beginner clips, per stroke class, against that stroke's own expert template.

stroke expert mean beginner mean experts higher?
backhand 88.9 86.5 yes
backhand_volley 89.5 85.8 yes
forehand 88.5 86.7 yes
forehand_volley 87.2 79.8 yes
serve 87.3 87.0 yes
smash 87.9 87.9 no

(Table above uses the current 7-joint angle set, Section 2.8 — regenerated after the shoulder joints were added; see Section 3.4 for what changed and the honest discussion of the smash row.) Experts scored higher on 5 of 6 strokes with held-out data (ml/skilleye/smoke_check_quality_scoring.py). This is a sanity check, not a validation -- it confirms the scoring system's direction is sane on the same subject-level proxy label used in Section 3.2, not that its absolute scores or flagged joints are correct at the level of a real coach's judgment. That remains gated on Path A/B ground truth (Section 5).

The aggregate numbers above are a directional average -- the actual per-clip output looks like this (two real held-out forehand_volley clips, the largest expert/beginner gap in the table above):

Quality-scoring example on two real held-out forehand_volley clips

The beginner clip's flagged joints are concrete and actionable (left elbow over-extended, right elbow under-extended, left knee not bent enough at contact) rather than an opaque number -- this is what "rule-based correction suggestions" (proposal Section 4.7) actually produces today, not a mockup of what it might eventually say.

3.4 Module A/B Validation

This section reports what happened when Module A (Section 2.8) and Module B (Section 2.9) were actually run and checked, including a result that did not come out clean — included here in full rather than left out, in keeping with how every other result in this document is reported.

Module A, same smoke check as Section 3.3, but with the new correlated scoring:

stroke expert mean beginner mean experts higher?
backhand 89.3 86.0 yes
backhand_volley 88.6 83.9 yes
forehand 86.2 83.2 yes
forehand_volley 87.1 77.0 yes
serve 87.6 85.5 yes
smash 86.7 87.0 no

(ml/skilleye/smoke_check_correlated_quality_scoring.py.) Same result as the independent scorer above: 5 of 6 strokes pass, smash does not. Because both the old (independent) and new (correlated) scorers fail on smash once the shoulder joints are included, this is not a bug introduced by Module A specifically -- it appears to be an effect of adding the shoulder angle, on a stroke class that only has 55-57 expert training clips (one of the thinnest classes). No attempt was made to adjust constants until the number looked better; that would be fitting the system to a desired answer rather than measuring it honestly. This is flagged as a real, open finding for further investigation (candidates: check whether beginners' shoulder angle happens to be unusually consistent for smash specifically, or whether 55-57 clips is simply too few for a stable 7-dimensional covariance estimate), not something papered over.

Module B, real hardware: the racket sensor was connected live and used to record a real 58-second volley take (Section 2.10). check_volley_swing_effort() was run against 28 segmented swings from that recording: 18 fell on the lower-force side, 10 on the higher-force side of the reference midpoint. Because that recording had no ball contact, this result is a pipeline check (the full sensor-to-rule chain works), not a skill assessment -- see Section 2.10 for why a ball-less swing isn't comparable to the paper's reference values.

3.5 Real-Recording Results

The six real takes from Section 2.11, scored end to end:

Stroke Score (v1) Score (Module A) Peak accel Module B / volley-effort flags
Backhand 83/100 81/100 13.2 g (130 m/s²) --
Forehand 89/100 86/100 12.6 g (123 m/s²) --
Backhand Volley 82/100 86/100 12.3 g (121 m/s²) 3 Module B flags; volley-effort: amateur-leaning
Forehand Volley 80/100 68/100 11.8 g (116 m/s²) volley-effort: amateur-leaning
Serve 87/100 85/100 17.5 g (172 m/s²) --
Smash 86/100 86/100 12.8 g (126 m/s²) --

The backhand volley take is the most interesting result: it's flagged by every Module B/effort check built in this project, on real self-recorded data, not a THETIS clip.

RTMPose skeleton at the contact frame, real backhand-volley take Real IMU trace for the same take -- accelerometer and gyroscope magnitude over time

Katsumi et al. (2026)'s two rules both fire: the shoulder-pelvis twist reverses sign between backswing and contact, and pelvic rotation is on the higher side at both phases -- the exact pattern the paper associates with less-skilled backhand volleys (§2.9). Aydin & Aydemir (2026)'s volley-effort check also flags this take (peak 121 m/s² against a 52.6 m/s² reference midpoint) -- and independently, the forehand-volley take does too (116 m/s²). Whether these particular takes had real ball contact wasn't recorded as metadata for this batch (unlike the explicitly ball-less test in Section 3.4) -- the peak values are real either way, but the comparison to Aydin & Aydemir's ball-contact reference should be read with that context marked unknown, not assumed favorable.

The other four strokes score in a broadly similar 80-89/100 range under both scorers, with no rule-based flags -- consistent with (though not proof of) those takes not containing the same patterns. Figures and raw numbers for all six takes: hardware/client/newresult_eval/ (results.json + one skeleton/IMU figure pair per stroke).

4. Discussion

4.1 Limitations

  • Camera geometry limits stroke disambiguation. The volley/groundstroke confusion in Section 3.1 is a direct consequence of THETIS's near-frontal camera — footwork and court-position cues that would resolve it are not visible from that angle. Side-view footage (planned, Section 5) should address this directly, not just improve accuracy but make the representation richer for the eventual per-phase error detection this project needs.
  • The beginner/expert label is a subject-level proxy, not a per-swing quality score. THETIS labels an entire subject as beginner or expert; it says nothing about whether a specific swing within that subject's clips was a particularly good or bad execution. The precision/recall asymmetry noted in Section 3.2 (the model calls some beginner clips "expert" and vice versa) is consistent with this — a beginner's better attempts plausibly look expert-ish in isolation, and the reverse. Section 3.2's result demonstrates that quality-relevant signal exists and is learnable, which de-risks the project's core premise; it does not by itself produce the continuous, per-swing quality score the finished system needs.
  • Frame rate. ~17–19fps may under-resolve the fastest phases of a serve or smash relative to the 60fps the original proposal assumed necessary; this has not yet been quantified directly (e.g. by comparing model behavior on down-sampled higher-fps footage) and is flagged as an open question rather than a settled limitation.
  • Sample size for cross-validation. 55 subjects is enough to show fold-to-fold variance (Section 3, 4.9 and 3.8 percentage-point standard deviations) but not enough to make that variance estimate itself highly precise — it indicates real single-split fragility rather than pinning down an exact confidence interval.

4.2 Implications for the Quality-Score System

Taken together, Sections 3.1 and 3.2 support two concrete design decisions for the system described in the proposal's Section 4.7: (1) stroke-type classification is solved to a practical standard on this representation and can run as a preprocessing step before stroke-specific quality rules are applied, and (2) motion quality is measurable from the same representation in principle, but the production quality-score model will need either per-swing ground truth (real coach ratings, or synthetic perturbation of known-good motion — both already planned as Path A/B) or a training signal richer than a subject-level label, because Section 3.2's result is a proof of learnability, not a finished quality metric. Section 2.6/3.3's rule-based scorer is a further step in that direction -- a working, demonstrable system rather than only a proof of concept -- but it is calibrated by a sanity check, not by the coach ratings or synthetic ground truth that would let its scores be trusted at face value.

5. Conclusion and Future Work

This work establishes, with real trained and cross-validated results rather than projected targets, that (1) a resumable, validated pipeline exists from raw racket-sport video to a normalized skeleton representation, (2) stroke type is classifiable from that representation to a practical standard (81.7% ± 4.9%, 6-way), and (3) motion quality signal — not just stroke identity — is present and learnable in the same representation (82.4% ± 3.8% on a weak proxy label), which supports the project's central premise ahead of collecting stronger ground truth. Sections 2.8-2.10 and 3.4 add to this: a correlated (not-independent) version of the quality scorer, two literature-grounded rules that specifically catch beginner mistakes, and a racket-mounted sensor that is now confirmed working on real hardware, not just designed on paper.

Remaining work, in priority order:

Done since the last pass (kept here briefly rather than silently dropped, so progress is visible): Module A and Module B are now wired into the demo UI (ml/skilleye/app.py) with a scoring-mode toggle and a skill-level-checks panel, not just reachable through their own scripts. 24 real synced webcam+IMU takes have been recorded across all 6 strokes (hardware/client/newresult/ + result/), and the full pipeline — capture, RTMPose, both scorers, Module B — has been run end to end on 6 of them (Section 2.11/3.5), the milestone item 6 below used to describe as not yet started.

  1. Team/university/advisor/contact fields in the competition proposal — still placeholders; blocks submission eligibility independent of all technical progress above.
  2. Path B: coach-rated ground truth — recruit coaches to blind-rate real swings, to be correlated against model output (target r > 0.7). Longest lead time item; should run in parallel with, not after, further technical work.
  3. Record ball-contact status as metadata, then re-check the volley-effort rule — 24 real volley takes now exist (Section 2.11/3.5 and the earlier batch), and two of them already flag check_volley_swing_effort() (Section 2.9) at real, high peak accelerations — but whether any of them had actual ball contact was never logged, so that comparison's context is currently unknown rather than confirmed. Cheap fix: extend sync_recorder_gui.py to ask before each take, then a real, trustworthy check is one recording session away.
  4. Investigate the smash smoke-check failure (Section 3.4) — both the independent and correlated scorers rate expert smash clips no higher than beginner ones once the shoulder joints are included, unlike every other stroke. Flagged honestly rather than silently adjusted away; needs a dedicated look (is it the shoulder joints specifically, or the smash class's small expert sample of 55-57 clips?) before smash's quality score can be trusted in a demo.
  5. Own side-view, higher-frame-rate recordings — required for the joint-angle-based error-detection rules in Section 4.7, and expected to resolve the volley/groundstroke confusion identified in Section 4.1. The sensor board itself is no longer the blocker here (Section 2.10/2.11); this is specifically about a side-on camera angle, which the current webcam recordings (frontal, like THETIS) don't provide either.
  6. A full camera+IMU recording session, then retrain the fusion model on real data — Section 2.11 proved the capture-through-evaluation pipeline works on real data, but that was one clip per stroke from one person; FusedBeginnerExpertModel (Section 2.7) still trains only on a synthetic signal derived from the skeleton, and retraining it on real data needs multiple people and multiple swings per stroke, following the same 5-fold rigor as Section 3.2. Protocol: docs/superpowers/specs/2026-07-23-imu-fusion-prototype-design.md.
  7. Address the THETIS/webcam domain gap (Section 2.11's first limit) — either by rebuilding the expert templates from the team's own recordings once enough exist, or by quantifying how much the camera/setup difference actually moves scores, so "domain gap, not calibrated" (currently qualitative) becomes a measured number.
  8. Path A: synthetic perturbation of known-good motion (known joint-angle/timing offsets injected into skilled-athlete clips) as a complementary, coach-independent ground-truth source for the quality-score model.
  9. A learned quality-score model (proposal Section 4.7) — Section 2.6/3.3/3.4 delivers a working rule-based v1 (phase detection, joint angles, expert-template comparison, now with a correlated-scoring option and two skill-level-specific rules); once ground truth from (2) and/or (8) is available, that data enables replacing or augmenting the rule-based scorer with a trained regression model, and calibrating FLAG_THRESHOLD/SCORE_SCALE against real quality judgments instead of a directional sanity check.
  10. Cross-dataset validation against the Wagner tennis serve dataset (Section 1.2) — its existing serve-quality labels are an external, independently-collected signal that could validate whether the quality-relevant patterns found in Section 3.2 generalize beyond THETIS and beyond amateur players. This requires extending the 2D pipeline to accept its 3D pose format (or projecting its 3D keypoints to 2D for direct compatibility) — not yet attempted, flagged here as a concrete, scoped next step rather than folded into the results above.

6. Reproducibility

ml/skilleye/                       pipeline and modeling code
  skeleton_pipeline.py             RTMPose output -> tracked subject -> normalized skeleton (§2.2)
  batch_extract.py                 CLI batch runner over a THETIS-shaped folder tree (resumable)
  skeleton_records.py              skeleton-JSON loading + subject-disjoint/k-fold splitting (torch-free)
  stroke_dataset.py                category merging (Table 1), torch Dataset -- re-exports skeleton_records.py
  stgcn_model.py                   ST-GCN architecture, COCO-17 skeleton graph (§2.3)
  train_stroke_classifier.py       stroke classifier, single split
  train_beginner_expert_stgcn.py   beginner/expert ST-GCN, single split (§2.4)
  beginner_expert_check.py         beginner/expert hand-crafted-feature baseline, v1 (§2.4)
  cross_validate.py                5-fold cross-validation for both classifiers (§2.5)
  generate_figures.py              renders ml/results/figures/*.png from the metrics JSONs
  quality/                         phase detection, joint angles, template scoring (§2.6)
  quality/correlation.py           conditional (correlated) z-score, Module A (§2.8)
  quality/skill_rules.py           skill-level-specific rules, Module B (§2.9)
  build_expert_templates.py        builds ml/results/quality_templates/templates.json (§2.6, §2.8)
  smoke_check_quality_scoring.py   experts-score-higher-than-beginners sanity check, independent scorer (§3.3)
  smoke_check_correlated_quality_scoring.py   same check, correlated scorer (§3.4)
  generate_qualitative_figures.py  renders the stroke-gallery and quality-comparison examples (§2.6/3.3)
  app.py                           Streamlit demo UI (§2.6) -- run: streamlit run app.py
  quality/llm_explainer.py         optional LLM-generated correction paragraphs (§2.6, needs NVIDIA_API_KEY)
  imu_fusion.py                    synthetic IMU signal + fusion model prototype (§2.7)
  train_beginner_expert_fusion_prototype.py   trains the fusion prototype (synthetic data, §2.7)
  requirements.txt
  README.md                        environment setup notes (incl. GPU-specific gotchas)

ml/results/
  RESULTS_SUMMARY.md                supplementary detail beyond what's inlined above
  figures/                          PNGs embedded above; regenerate with generate_figures.py / generate_qualitative_figures.py
  quality_templates/                templates.json: per-stroke expert (phase, joint) statistics + covariance (§2.6, §2.8)
  imu_fusion_prototype/             prototype-only metrics (synthetic IMU data, §2.7) -- not a benchmark
  cross_validation/                 §3 headline numbers: per-fold + aggregated metrics
  beginner_expert_check.json        v1 (§2.4): hand-crafted features + logistic regression
  beginner_expert_stgcn/            v2 (§2.4): ST-GCN, single split, trained weights + metrics
  stroke_classifier/                early 5-class iteration, kept for reference
  stroke_classifier_v2/             6-class + augmentation, single split, trained weights + metrics

hardware/skilleye_2.0/             KiCad rev2.1 PCB project (§2.7, finalized): GY-521/MPU6050 +
                                    Seeed XIAO + TP4056 charging -- built and streaming (§2.10)
hardware/skilleye_1.0/             earlier rev1.0/1.1 custom-board iteration, kept for reference
hardware/gerbers/                  fabrication-ready gerber files for the rev2.1 board
hardware/firmware/firmware.ino     rev2.x streaming firmware, XIAO ESP32-C6 + GY-521 (§2.10)
hardware/client/imu_client.py      computer-side recording/monitoring client (§2.10)
hardware/client/live_buffer.py     rolling-window buffer for the live dashboard, unit-tested (§2.10)
hardware/client/live_dashboard.py  live IMU monitoring Streamlit page (§2.10)
hardware/client/sync_recorder.py   one-command webcam+IMU recorder, unit-tested (§2.10, §5 item 6)
hardware/client/sync_recorder_gui.py   Start/Stop + live preview + Save-As dialog version (§2.10)
hardware/client/requirements.txt   opencv-python, pillow -- only the two sync_recorder* tools need these
hardware/client/recorded/*.csv     real recordings: backhand/forehand/serve (with and without
                                    ball contact) and one ball-less volley take (§2.10)
hardware/client/newresult/         6 real synced webcam+IMU takes (1 per stroke), fully evaluated (§2.11);
                                    .mp4s kept locally only (large binary, see .gitignore, same as THETIS/)
hardware/client/result/            18 more real IMU+video takes (pre-alignment.json-fix batch); IMU
                                    CSVs and video timestamps tracked, .mp4s local-only as above
ml/skilleye/extract_real_recordings.py   RTMPose on the real recordings above (§2.11)
ml/skilleye/evaluate_real_recordings.py  scores + figures for the real recordings (§2.11, §3.5)
hardware/client/newresult_skeletons/     RTMPose output for the 6 evaluated takes (§2.11)
hardware/client/newresult_eval/          §3.5's figures + results.json
docs/schematics/2.0/rev2.1/        rendered schematic PDF/SVG for the current (rev2.1) board
docs/schematics/1.0/               rendered schematics for the earlier rev1.0/1.1 iteration
docs/layouts/rev2.0/               fabrication layout PDFs (front/back panel)
docs/schematics/system_architecture.svg          full pipeline diagram (§2.0), hand-authored
docs/schematics/sensor_*_concept.svg             placement/mechanism diagrams (§2.7) --
                                    illustrative, not photos, predate the built board
docs/superpowers/specs/2026-08-14-correlated-zscore-module-a-design.md    Module A design (§2.8)
docs/superpowers/specs/2026-08-14-skill-level-rules-module-b-design.md    Module B design (§2.9)
docs/superpowers/specs/2026-08-14-live-imu-dashboard-design.md            live dashboard design (§2.10)
docs/superpowers/specs/2026-08-20-synced-video-imu-recorder-design.md     sync_recorder.py design (§2.10)
docs/superpowers/specs/2026-08-20-sync-recorder-gui-design.md             sync_recorder_gui.py design (§2.10)
docs/superpowers/specs/2026-08-17-wire-modules-ab-into-demo-ui-design.md  app.py wiring design (§5 item 7)
skilleye-demo/                     bundled standalone demo (app.py + trained weights), same code as ml/skilleye/
skilleye-website/                  results-showcase webpage

Dataset (not included in this repository — see ml/skilleye/README.md for how to fetch it): THETIS, VIDEO_RGB + papers subsets only.

References

  • Gourgari, S., Goudelis, G., Karpouzis, K., & Kollias, S. (2013). THETIS: Three Dimensional Tennis Shots a Human Action Dataset. CVPR Workshops.
  • Yan, S., Xiong, Y., & Lin, D. (2018). Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. AAAI.
  • Jiang, T., et al. (2023). RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose. arXiv:2303.07399.
  • Wagner, J. (2024). Tennis Serve Analysis Dataset: 3D Pose Sequences from 2024 US Open Broadcast Video. https://github.com/jasnwag/tennis_serve_dataset (CC BY 4.0).
  • Elliott, B. (2006). Biomechanics and tennis. British Journal of Sports Medicine.
  • Knudson, D., & Elliott, B. (2004). Biomechanics of Tennis Strokes.
  • Katsumi, K., Koda, H., & Kida, N. (2026). Analysis of Upper-Limb Movement Characteristics in Tennis Volleys Based on Skill-Level Differences: Kinematic Features of the Backhand Versus Forehand Volley. Journal of Functional Morphology and Kinesiology, 11(2), 203.
  • Aydin, E.H., & Aydemir, O. (2026). A Robust Deep Learning Framework for Skill Level Discrimination in Tennis Strokes Using Bilateral IMU Measurements. Sensors, 26(10), 3273.

The Katsumi and Aydin & Aydemir figures cited in Sections 2.9 and 3.4 were extracted via an automated read of each paper's published page rather than a manual PDF read — worth a spot-check against the source before citing in a formal submission.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages