Multimodal Imagination for Reasoning Assessment, a test that measures Reasoning models need not only words but also images
🟠 What is the idea?
- Where text doesn't help, images sharply improve the model's thinking.
- If the model is given drawings of intermediate steps, accuracy on average increases by 33.7%.
- The benchmark includes 546 tasks in 20 categories where you need to "see" rather than just read: blocks, mirrors, trajectories, forces, etc.
How the test is structured?
- direct question
- reasoning with text
- reasoning with visual steps (sketches)
🟢 What was found?
- Text alone often worsens performance because words poorly describe space.
- If the model is given images, the result improves significantly, especially in the exact sciences.
In the benchmark: 546 tasks in geometry, physics, logical puzzles, and causal relations.
Testing modes:
• Direct - model answers directly
• Text-CoT - text chain-of-thought
• Visual-CoT - model reasons through drawings and visual steps
• No model exceeded 20% accuracy in Direct mode (GPT-5 ~16.5%)
• Text-CoT often worsens results (e.g., −18% for Gemini 2.5 Pro)
• Visual-CoT gives an average gain of +33.7%, especially noticeable in physics tasks
Models need a *visual way of thinking*.
They need to be able to read simple diagrams, understand them and use them in reasoning, otherwise many tasks simply remain unsolvable.
🤖 Data Science, ML & Big Data with @DataXplore
