Recognizing fine-grained actions from temporally corrupted skeleton sequences remains a significant challenge, particularly in real-world scenarios where online pose estimation often yields substantial missing data.
Existing methods often struggle to accurately recover temporal dynamics and fine-grained spatial structures, resulting in the loss of subtle motion cues crucial for distinguishing similar actions.
To address this, we propose FineTec, a unified framework for Fine-grained action recognition under Temporal Corruption.
FineTec first restores a base skeleton sequence from corrupted input using context-aware completion with diverse temporal masking.
Next, a skeleton-based spatial decomposition module partitions the skeleton into five semantic regions, further divides them into dynamic and static subgroups based on motion variance, and generates two augmented skeleton sequences via targeted perturbation.
These, along with the base sequence, are then processed by a physics-driven estimation module, which utilizes Lagrangian dynamics to estimate joint accelerations.
Finally, both the fused skeleton position sequence and the fused acceleration sequence are jointly fed into a GCN-based action recognition head.
Extensive experiments on both coarse-grained (NTU-60, NTU-120) and fine-grained (Gym99, Gym288) benchmarks show that FineTec significantly outperforms previous methods under various levels of temporal corruption.
Specifically, FineTec achieves top-1 accuracies of 89.1% and 78.1% on the challenging Gym99-severe and Gym288-severe settings, respectively, demonstrating its robustness and generalizability.
Method
FineTec consists of three core modules:
â‘ Context-aware Sequence Completion restores missing or corrupted skeleton frames using in-context learning, producing $S_{base}$;
② Skeleton-based Spatial Decomposition partitions $S_{base}$ into anatomical regions by motion intensity, generating dynamic ($S_{dyna}$) and static ($S_{stat}$) variants, which are fused into $S_{pred}$;
③ Physics-driven Acceleration Modeling infers joint accelerations via Lagrangian dynamics and data-driven finite differences, producing fused temporal dynamics features $\mathbf{a}$.
The resulting positional ($S_{pred}$) and dynamic ($\mathbf{a}_{pred}$) features are used for downstream fine-grained action recognition.
For more details, please refer to our paper.
Results
The main quantitative results on the two fine-grained skeleton datasets, Gym99-skeleton and Gym288-skeleton, are presented in the table above.
These results are reported across three difficulty levels: minor (25% frame missing), moderate (50% frame missing), and severe (75% frame missing).
It can be observed that the proposed FineTec framework consistently achieves the best performance under all conditions.
Notably, in the most challenging scenario—Gym288-skeleton with severe frame missing—FineTec attains a Top-1 accuracy of 78.1%, surpassing all previous skeleton-based methods.
In terms of mean class accuracy, FineTec improves upon the best baseline by 13%, and outperforms the latest work by 50%.
Overall, these results demonstrate that FineTec achieves outstanding effectiveness across fine-grained datasets and under all levels of difficulty.
For additional results and ablation studies, please refer to our paper.
Demo Video
Dataset
Overview
FineGym-skeleton is a human skeleton-based action recognition benchmark derived from the Gym99 and Gym288 subsets of the
FineGym dataset.
It provides temporally precise, fine-grained annotations of gymnastic actions together with 2D human pose sequences extracted from RGB subaction clips. The V2 release contains Gym99-skeleton-V2 and Gym288-skeleton-V2.
This dataset is designed to support research in:
Fine-grained action recognition
Temporally corrupted or incomplete action modeling
Skeleton-based representation learning
Physics-aware motion understanding
Key Statistics
Item
Gym99-skeleton
Gym288-skeleton
Action classes
99
288
Total instances
34,803
38,935
Training samples
26,282
29,290
Validation samples
8,521
9,645
Total annotated frames
1,617,291
1,882,226
Frames per sample (min / mean / max)
2 / 46.47 / 725
2 / 48.34 / 725
Each sample contains one tracked gymnast represented by 17 COCO-style 2D keypoints per frame. The action classes cover Floor Exercise (FX), Balance Beam (BB), Uneven Bars (UB), and Vault – Women (VT).
Annotation Pipeline
For each RGB subaction clip, a bounding box was manually annotated on the first frame to identify the target gymnast. OSTrack was then used to track the target throughout the clip, followed by HRNet for frame-by-frame skeleton keypoint extraction.
Dataset Structure
Each skeleton subset is distributed as a Python dictionary with two top-level keys: split and annotations.
Top-Level Keys:
• split: Dictionary containing train and val sample-ID lists
â—¦ Gym99-skeleton: 26,282 training IDs and 8,521 validation IDs
â—¦ Gym288-skeleton: 29,290 training IDs and 9,645 validation IDs
• annotations: One annotation dictionary per action instance
Annotation Fields:
Key
Type
Shape / Example
Description
frame_dir
str
"A0xAXXysHUo_002184_002237_0035_0036"
Unique identifier for the action clip
label
int
93
Zero-based class label (0–98 for Gym99 or 0–287 for Gym288)
img_shape
tuple
(720, 1280)
Height and width of original video frames
original_shape
tuple
(720, 1280)
Same as img_shape (for compatibility)
total_frames
int
48
Number of frames in the action sequence
keypoint
np.ndarray (float16)
(1, T, 17, 2)
2D joint coordinates (x, y) for 17 COCO keypoints over T frames
keypoint_score
np.ndarray (float16)
(1, T, 17)
Per-keypoint pose-estimation scores
Note: The first dimension (1) in keypoint and keypoint_score corresponds to the number of persons (always 1 in this dataset).
Action Classes
Gym99-skeleton contains 99 classes, while Gym288-skeleton contains 288 classes. Both subsets cover four apparatuses: Floor Exercise (FX), Balance Beam (BB), Uneven Bars (UB), and Vault – Women (VT).
Each class represents a highly specific movement (e.g., "Switch leap with 0.5 turn", "Clear hip circle backward with 1 turn to handstand"),
reflecting the fine-grained nature of competitive gymnastics scoring.
The FineGym-RGB-subactions release contains 39,092 MP4 clips, organized into four parts with a total size of approximately 8.39 GB.
Directory
MP4 files
Size
part1
10,000
2.04 GB
part2
10,000
2.20 GB
part3
10,000
2.19 GB
part4
9,092
1.95 GB
Total
39,092
8.39 GB
An example filename is 0LtLS9wROrk_E_000147_000152_A_0000_0005.mp4, which preserves the source-video ID and the zero-padded temporal identifiers used by the FineGym event/action annotations.
Download & License
The skeleton annotations and RGB subaction clips are available on Hugging Face
under the CC-BY-4.0 license.
Acknowledgements
We thank the authors of FineGym
for their foundational work in fine-grained action recognition. Target gymnasts were tracked with OSTrack, and skeleton keypoints were extracted with HRNet. If you use this dataset, please cite both FineTec and the original FineGym paper.
Citation
@article{shao2026finetec,
title={FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion},
volume={40},
url={https://ojs.aaai.org/index.php/AAAI/article/view/37838},
DOI={10.1609/aaai.v40i11.37838},
number={11},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
author={Shao, Dian and Shi, Mingfei and Liu, Like},
year={2026},
month={Mar.},
pages={8842-8850}
}