Skip to content

Submitting Results for DeSTA2.5-Audio #8

Description

@kehanlu

Dear authors,

We have inferenced DeSTA2.5-Audio on MMAR. We would like to share these results to help update the leaderboard and provide a reference for the community.

DeSTA2.5-Audio: https://arxiv.org/abs/2507.02768
Github: https://github.com/kehanlu/DeSTA2.5-Audio

We prompt model with direct answer:

"messages": [
      {
        "role": "system",
        "content": "Focus on the audio clip and instruction. Output your answer in the format \"The correct answer is: ___\"."
      },
      {
        "role": "user",
        "content": "<|AUDIO|>\n\nDetermine what is producing the sound in the audio Choose from the following options: \"Owl\", \"Robot\", \"Rooster\" or \"Parrot\""
      }
    ],
******************************
Modality-wise Accuracy:
sound : 38.18% over 165 samples
music : 40.78% over 206 samples
speech : 59.18% over 294 samples
mix-sound-music : 54.55% over 11 samples
mix-sound-speech : 57.34% over 218 samples
mix-music-speech : 58.54% over 82 samples
mix-sound-music-speech : 33.33% over 24 samples
******************************
Category-wise Accuracy:
Signal Layer : 55.81% over 43 samples
Perception Layer : 41.83% over 404 samples
Semantic Layer : 59.22% over 412 samples
Cultural Layer : 50.35% over 141 samples
******************************
Sub-category-wise Accuracy:
Speaker Analysis : 60.42% over 48 samples
Environmental Perception and Reasoning : 53.69% over 149 samples
Content Analysis : 59.21% over 304 samples
Correlation Analysis : 46.00% over 50 samples
Counting and Statistics : 27.27% over 99 samples
Professional Knowledge and Reasoning : 47.89% over 71 samples
Culture of Speaker : 59.62% over 52 samples
Aesthetic Evaluation : 25.00% over 8 samples
Emotion and Intention : 58.33% over 60 samples
Anomaly Detection : 82.35% over 17 samples
Spatial Analysis : 46.67% over 15 samples
Temporal Analysis : 21.43% over 28 samples
Acoustic Quality Analysis : 44.44% over 18 samples
Music Theory : 41.27% over 63 samples
Audio Difference Analysis : 25.00% over 8 samples
Imagination : 40.00% over 10 samples
******************************
Total Accuracy: 50.80% over 1000 samples
******************************
No prediction count: 0

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions