A benchmark for multimodal retrieval
ARK-Bench
Knowledge- and Reasoning-Intensive Multimodal Retrieval
Introduction
Multimodal retrieval often relies on more than matching a query to visually or semantically similar content. Finding the correct candidate can require specialized knowledge, understanding spatial relationships, or reasoning across multiple pieces of evidence. Existing benchmarks largely emphasize semantic matching on everyday images, offering limited insight into these capabilities.
ARK-Bench evaluates retrieval from two complementary perspectives: the knowledge domains a task draws on and the reasoning skills needed to identify the correct candidate. It spans five knowledge domains with 17 subtypes and six reasoning skills, covering 16 visual data types with both unimodal and multimodal queries and candidates. Most queries include targeted hard negatives to reduce shortcut matching. Instances are also categorized by cognitive demand: perception-only, knowledge-intensive, reasoning-intensive, or both knowledge- and reasoning-intensive.
This leaderboard focuses on embedding models. Start with the Overview to compare overall performance, then explore individual knowledge domains, reasoning skills, and cognitive demands to examine each model’s strengths and limitations.
Explore the results
Leaderboard
| Loading results… |
Contribute to ARK-Bench
Submit results
Upload topic-level retrieval results for automatic scoring.
Reviewed submissions can then be added to the leaderboard.
Prepare predictions
Rank each query against its topic’s full Gallery.
Fill 20retrieved_idsfrom highest to lowest relevance.Click to preview a prediction row
{"topic":"Example","query_id":0,"retrieved_ids":[17,4,...,5,2]}Evaluate your model
Upload predictions to calculate aggregate scores.
Submission report
Review or download the scores, then submit the report.
Await maintainer review
Approved results are added to the leaderboard.
Evaluate predictions
Upload your model’s results
Preparing upload…