It is a re-implementation code for the SGST model.
- Kao Zhang, Tao Song, Zhihua Hu, Ming Li, Xin Ding. SGST-Transformer: A Spherical Geometry-Aware Spatio-Temporal Transformer for 360 Video Saliency Prediction. IEEE CVPR Findings 2026.
Paper,
Supplemental,
Poster.
Github: https://github.com/zhangkao/3DIP-SGST
And it is easy to change the output format in our code.
- The results of video task is saved by ".mat"(uint8) formats.
- You can get the color visualization results based on the "Visualization Tools".
- You can evaluate the performance based on the "EvalScores Tools".
Results: ALL (4.38G): Sports360 (2.6G), SVGC-AVA (1.5G), AVS-ODV (224M), 360AV-HM (42M)
This research was funded by: the National Natural Science Foundation of China (Grant No. 62201404), the Startup Foundation for Introducing Talent of NUIST (Grant No. 2024r061), and the Postgraduate Research & Practice Innovation Program of Jiangsu Province (Grant No. KYCX25_1654).
If you use the SGST 360° video saliency model, please cite the following paper:
@InProceedings{Zhang_2026_CVPR,
author = {Zhang, Kao and Song, Tao and Hu, Zhihua and Li, Ming and Ding, Xin},
title = {SGST-Transformer: A Spherical Geometry-Aware Spatio-Temporal Transformer for 360deg Video Saliency Prediction},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings},
month = {June},
year = {2026},
pages = {2596-2605}
}
Kao ZHANG
3D Reconstruction and Image Processing Group (3DIP)
Perceptual and Generative AI Lab (PGAI Lab)
Nanjing University of Information Science and Technology, Nanjing, China.
Email: kaozhang@nuist.edu.cn
Tao SONG
3D Reconstruction and Image Processing Group (3DIP)
Perceptual and Generative AI Lab (PGAI Lab)
Nanjing University of Information Science and Technology, Nanjing, China.
Email: taosong@nuist.edu.cn