STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition
STAG is a new research model for micro-expression recognition that tries to fix a common weakness in prior approaches: over-reliance on apex-onset frames and weak modeling of the tiny motion changes between them. The paper frames the problem as one of both spatial reasoning and temporal dynamics, then tackles it with a dual-branch network that processes facial structure and motion together.
The spatial side uses an enhanced graph attention network, while the temporal side uses a transformer encoder. A bidirectional cross-attention module lets the two branches refine each other, and the model also introduces AU-guided dynamic connectivity so facial regions can connect differently depending on which action units appear active. In practice, that means the network is trying to learn a more expression-aware facial graph rather than a fixed one.
The authors say they extract optical flow from discriminative frames using magnitude-based selection and temporal attention, then optimize the fused representation with focal loss. They evaluate on CASME II, 4DME, DFME, NaME, SAMM, and SMIC-HS, and report better robustness, generalization, interpretability, and computational efficiency than prior methods.
For game developers, this is mostly relevant if you work on facial analysis, performance capture tooling, emotion-aware NPCs, or any pipeline that needs to read very subtle facial motion. The broader takeaway is that hybrid graph-plus-transformer designs are still proving useful when the signal is sparse and timing matters, especially when you need cross-dataset generalization...
- what
- STAG proposes a dynamic ROI-AU-coupled spatial-temporal network for micro-expression recognition.
- who
- Authors: Nandani Sharma, Varun Sharma, and Dinesh Singh.
- when
- Submitted to arXiv on 26 Jun 2026 (arXiv:2606.28083).
- impact
- Could inform facial-analysis, performance-capture, and emotion-reading systems that need to detect subtle facial motion robustly.
Promising method for subtle facial-motion analysis
Follow computer-vision updates
See relevant stories in your personalized news feed.
Discussion