A benchmark for multimodal retrieval

ARK-Bench

Knowledge- and Reasoning-Intensive Multimodal Retrieval

Introduction

Multimodal retrieval often relies on more than matching a query to visually or semantically similar content. Finding the correct candidate can require specialized knowledge, understanding spatial relationships, or reasoning across multiple pieces of evidence. Existing benchmarks largely emphasize semantic matching on everyday images, offering limited insight into these capabilities.

ARK-Bench evaluates retrieval from two complementary perspectives: the knowledge domains a task draws on and the reasoning skills needed to identify the correct candidate. It spans five knowledge domains with 17 subtypes and six reasoning skills, covering 16 visual data types with both unimodal and multimodal queries and candidates. Most queries include targeted hard negatives to reduce shortcut matching. Instances are also categorized by cognitive demand: perception-only, knowledge-intensive, reasoning-intensive, or both knowledge- and reasoning-intensive.

This leaderboard focuses on embedding models. Start with the Overview to compare overall performance, then explore individual knowledge domains, reasoning skills, and cognitive demands to examine each model’s strengths and limitations.

Explore the results

Leaderboard

ARK-Bench model results
Loading results…

Contribute to ARK-Bench

Submit results

Upload topic-level retrieval results for automatic scoring.
Reviewed submissions can then be added to the leaderboard.

  1. Prepare predictions

    Rank each query against its topic’s full Gallery.
    Fill 20 retrieved_ids from highest to lowest relevance.

    Click to preview a prediction row {"topic":"Example","query_id":0,"retrieved_ids":[17,4,...,5,2]}
  2. Evaluate your model

    Upload predictions to calculate aggregate scores.

  3. Submission report

    Review or download the scores, then submit the report.

  4. Await maintainer review

    Approved results are added to the leaderboard.

Evaluate predictions

Upload your model’s results

Preparing upload…

Submission details One JSON object per line. Maximum upload: 2 MiB.

Report submitted

Awaiting maintainer review

Your evaluation report has been submitted successfully. A maintainer will review it before publication.