FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion

AAAI 2026

Dian Shao†, Mingfei Shi, Like Liu

Northwestern Polytechnical University
†Corresponding Author

📄 AAAI Paper 📝 arXiv 💻 Code 💾 Dataset 📋 BibTeX

Note: Supplementary materials are available in the arXiv version.

Abstract

FineTec Poster

Poster

🔍 View Full Size

Recognizing fine-grained actions from temporally corrupted skeleton sequences remains a significant challenge, particularly in real-world scenarios where online pose estimation often yields substantial missing data. Existing methods often struggle to accurately recover temporal dynamics and fine-grained spatial structures, resulting in the loss of subtle motion cues crucial for distinguishing similar actions. To address this, we propose FineTec, a unified framework for Fine-grained action recognition under Temporal Corruption. FineTec first restores a base skeleton sequence from corrupted input using context-aware completion with diverse temporal masking. Next, a skeleton-based spatial decomposition module partitions the skeleton into five semantic regions, further divides them into dynamic and static subgroups based on motion variance, and generates two augmented skeleton sequences via targeted perturbation. These, along with the base sequence, are then processed by a physics-driven estimation module, which utilizes Lagrangian dynamics to estimate joint accelerations. Finally, both the fused skeleton position sequence and the fused acceleration sequence are jointly fed into a GCN-based action recognition head. Extensive experiments on both coarse-grained (NTU-60, NTU-120) and fine-grained (Gym99, Gym288) benchmarks show that FineTec significantly outperforms previous methods under various levels of temporal corruption. Specifically, FineTec achieves top-1 accuracies of 89.1% and 78.1% on the challenging Gym99-severe and Gym288-severe settings, respectively, demonstrating its robustness and generalizability.

Method

FineTec Architecture

FineTec consists of three core modules: ① Context-aware Sequence Completion restores missing or corrupted skeleton frames using in-context learning, producing $S_{base}$; ② Skeleton-based Spatial Decomposition partitions $S_{base}$ into anatomical regions by motion intensity, generating dynamic ($S_{dyna}$) and static ($S_{stat}$) variants, which are fused into $S_{pred}$; ③ Physics-driven Acceleration Modeling infers joint accelerations via Lagrangian dynamics and data-driven finite differences, producing fused temporal dynamics features $\mathbf{a}$. The resulting positional ($S_{pred}$) and dynamic ($\mathbf{a}_{pred}$) features are used for downstream fine-grained action recognition. For more details, please refer to our paper.

Results

FineTec Results

The main quantitative results on the two fine-grained skeleton datasets, Gym99-skeleton and Gym288-skeleton, are presented in the table above. These results are reported across three difficulty levels: minor (25% frame missing), moderate (50% frame missing), and severe (75% frame missing). It can be observed that the proposed FineTec framework consistently achieves the best performance under all conditions. Notably, in the most challenging scenario—Gym288-skeleton with severe frame missing—FineTec attains a Top-1 accuracy of 78.1%, surpassing all previous skeleton-based methods. In terms of mean class accuracy, FineTec improves upon the best baseline by 13%, and outperforms the latest work by 50%. Overall, these results demonstrate that FineTec achieves outstanding effectiveness across fine-grained datasets and under all levels of difficulty. For additional results and ablation studies, please refer to our paper.

Demo Video

Dataset

Gym288-skeleton (V1) Dataset Visualization

Overview

FineGym-skeleton is a human skeleton-based action recognition benchmark derived from the Gym99 and Gym288 subsets of the FineGym dataset. It provides temporally precise, fine-grained annotations of gymnastic actions together with 2D human pose sequences extracted from RGB subaction clips. The V2 release contains Gym99-skeleton-V2 and Gym288-skeleton-V2.

This dataset is designed to support research in:

Key Statistics

Item Gym99-skeleton Gym288-skeleton
Action classes99288
Total instances34,80338,935
Training samples26,28229,290
Validation samples8,5219,645
Total annotated frames1,617,2911,882,226
Frames per sample (min / mean / max)2 / 46.47 / 7252 / 48.34 / 725

Each sample contains one tracked gymnast represented by 17 COCO-style 2D keypoints per frame. The action classes cover Floor Exercise (FX), Balance Beam (BB), Uneven Bars (UB), and Vault – Women (VT).

Annotation Pipeline

For each RGB subaction clip, a bounding box was manually annotated on the first frame to identify the target gymnast. OSTrack was then used to track the target throughout the clip, followed by HRNet for frame-by-frame skeleton keypoint extraction.

Dataset Structure

Each skeleton subset is distributed as a Python dictionary with two top-level keys: split and annotations.

Top-Level Keys:

• split: Dictionary containing train and val sample-ID lists

â—¦ Gym99-skeleton: 26,282 training IDs and 8,521 validation IDs

â—¦ Gym288-skeleton: 29,290 training IDs and 9,645 validation IDs

• annotations: One annotation dictionary per action instance

Annotation Fields:

Key Type Shape / Example Description
frame_dir str "A0xAXXysHUo_002184_002237_0035_0036" Unique identifier for the action clip
label int 93 Zero-based class label (0–98 for Gym99 or 0–287 for Gym288)
img_shape tuple (720, 1280) Height and width of original video frames
original_shape tuple (720, 1280) Same as img_shape (for compatibility)
total_frames int 48 Number of frames in the action sequence
keypoint np.ndarray (float16) (1, T, 17, 2) 2D joint coordinates (x, y) for 17 COCO keypoints over T frames
keypoint_score np.ndarray (float16) (1, T, 17) Per-keypoint pose-estimation scores

Note: The first dimension (1) in keypoint and keypoint_score corresponds to the number of persons (always 1 in this dataset).

Action Classes

Gym99-skeleton contains 99 classes, while Gym288-skeleton contains 288 classes. Both subsets cover four apparatuses: Floor Exercise (FX), Balance Beam (BB), Uneven Bars (UB), and Vault – Women (VT). Each class represents a highly specific movement (e.g., "Switch leap with 0.5 turn", "Clear hip circle backward with 1 turn to handstand"), reflecting the fine-grained nature of competitive gymnastics scoring.

For the full list of class names and mappings, please refer to the FineGym website and the original CVPR 2020 paper.

Usage Example

import pickle

# Load the dataset
with open("Gym288-skeleton-V2.pkl", "rb") as f:
    data = pickle.load(f)

# Access training samples
train_ids = data["split"]["train"]  # 29,290 samples
val_ids = data["split"]["val"]      # 9,645 samples

# Access annotations
sample = data["annotations"][0]
print("Label:", sample["label"])
print("Frames:", sample["total_frames"])
print("Keypoints shape:", sample["keypoint"].shape)  # (1, T, 17, 2)

# Extract skeleton sequence
skeleton_seq = sample["keypoint"][0]  # (T, 17, 2)

FineGym-RGB-subactions

The FineGym-RGB-subactions release contains 39,092 MP4 clips, organized into four parts with a total size of approximately 8.39 GB.

Directory MP4 files Size
part110,0002.04 GB
part210,0002.20 GB
part310,0002.19 GB
part49,0921.95 GB
Total39,0928.39 GB

An example filename is 0LtLS9wROrk_E_000147_000152_A_0000_0005.mp4, which preserves the source-video ID and the zero-padded temporal identifiers used by the FineGym event/action annotations.

Download & License

The skeleton annotations and RGB subaction clips are available on Hugging Face under the CC-BY-4.0 license.

Acknowledgements

We thank the authors of FineGym for their foundational work in fine-grained action recognition. Target gymnasts were tracked with OSTrack, and skeleton keypoints were extracted with HRNet. If you use this dataset, please cite both FineTec and the original FineGym paper.

Citation

@article{shao2026finetec,
  title={FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion},
  volume={40},
  url={https://ojs.aaai.org/index.php/AAAI/article/view/37838},
  DOI={10.1609/aaai.v40i11.37838},
  number={11},
  journal={Proceedings of the AAAI Conference on Artificial Intelligence},
  author={Shao, Dian and Shi, Mingfei and Liu, Like},
  year={2026},
  month={Mar.},
  pages={8842-8850}
}

Contact

For questions about this work, please contact:
mingfeishi5@mail.nwpu.edu.cn

×