Can search agents answer complex questions posed in any language?
MAST is a new shared track at FIRE 2026 that evaluates the performance of multilingual agentic systems (multilingual search & agents) to retrieve evidence & produce correct answers for queries posed on 21 typologically diverse languages.
A cross-lingual agentic information retrieval challenge
Agents combined with search tools has demonstrated impressive capabilities on complex fact-seeking benchmarks such as BrowseComp [1]. However, existing search agent benchmarks are mostly in English — leaving the question unanswered: how well do agentic search systems perform when the user's queries are expressed in their potentially native and non-English languages?
MAST benchmark is focused on evaluation of multilingual agentic search systems. In our inaugral 2026 year, we are building on BrowseComp-Plus (ACL 2026) [2], a reproducible & verifiable extension of BrowseComp with 830 challenging English queries with a verified English corpus of ~100,000 web-sourced documents and human judgments. This year, we are translating queries across 15 typologically diverse languages potentially spoken by over 1.6 billion people.
This year's task is cross-lingual: queries are asked in one language, documents & answers are in English, and systems can reason, iteratively search, plan, and synthesize in either language. e.g., a search agent can ask a query in Chinese, use multilingual retriever to retrieve English documents for a given query in Chinese.
Next year, based on the feedback from this year's participants, we plan to extend MAST to use a multilingual corpus and human-written queries in 15+ languages.
Choose the track that fits your research interest
Cross-lingual agentic search in 15 typologically diverse languages
Participating systems receive queries in one of 15 languages spanning high, mid, and low-resource settings. The goal is to retrieve the relevant documents from BrowseComp-Plus corpus and produce a correct, concise answer in English.
Deep-research agents for the Indian languages
A dedicated track for Indic languages. MAST Indic focuses on 9 Indic languages, morphologically rich languages, and spoken commonly in India.
Track 1 — MAST Multilingual: 15 typologically diverse languages
| Language | Code | Family | Regions of Speakers | Script | Speakers (M) | Resource Tier |
|---|---|---|---|---|---|---|
| 🟢 High Resource | ||||||
| English | en | Indo-European | United States, United Kingdom, Canada, Australia and globally | Latin | 1,500+ | High |
| Chinese | zh | Sino-Tibetan | China, Taiwan, Singapore, Malaysia | Han | 1,351 | High |
| French | fr | Indo-European | France, DR Congo, Canada, Belgium, Switzerland | Latin | 332 | High |
| Russian | ru | Indo-European | Russia, Belarus, Kazakhstan, Kyrgyzstan, Tajikistan | Cyrillic | 199 | High |
| Spanish | es | Indo-European | Spain, Mexico, Latin America | Latin | 510 | High |
| German | de | Indo-European | Germany, Austria, Switzerland, Liechtenstein | Latin | 133 | High |
| 🟡 Mid Resource | ||||||
| Arabic | ar | Afro-Asiatic | Egypt, Algeria, Saudi Arabia, Iraq, Morocco | Arabic | 375 | Mid |
| Bengali | bn | Indo-European | Bangladesh, India | Bengali | 280 | Mid |
| Finnish | fi | Uralic | Finland, Sweden | Latin | 6 | Mid |
| Hindi | hi | Indo-European | India | Devanagari | 584 | Mid |
| Thai | th | Kra-Dai | Thailand | Thai | 56 | Mid |
| Urdu | ur | Indo-European | Pakistan, India | Arabic | 313 | Mid |
| Tamil | ta | Dravidian | India, Sri Lanka, Singapore, Malaysia | Tamil | 91 | Mid |
| 🔴 Low Resource | ||||||
| Swahili | sw | Niger-Congo | Tanzania, Kenya, Uganda, DR Congo | Latin | 194 | Low |
| Welsh | cy | Indo-European | Wales, Argentina | Latin | 0.92 | Low |
What systems must do, and how they are measured
Given a query in a non-English language and the full BrowseComp-Plus corpus, participating systems must retrieve relevant evidence documents In English and produce a short and correct answer in English. Systems can be fully agentic involving open-source or closed-source systems and tools.
Every team must submit a list of retrieved document IDs per search turn. We will evaluate Recall against the human-verified relevance labels from BrowseComp-Plus, computed per language and aggregated per resource tier. We also evaluate search efficiency, i.e., the number of search turns required to reach the correct answer.
The headline leaderboard metric is Exact Match (EM) accuracy: the system's predicted English Exact Answer is compared against the BrowseComp-Plus ground-truth answer after normalisation, and averaged over languages. An LLM judge — either open-source (e.g., Qwen3-32B) or closed-source (e.g., Opus 5) — is additionally used to adjudicate semantically equivalent answers.
Runs are submitted as JSONL files containing reasoning steps, tool calls, retrieved document IDs per search round, and a final output answer. Maximum 3 runs per language per team.
View format spec →Official ranking by Exact Match (EM) accuracy on the held-out test queries
Track 1 — MAST Multilingual: 15 typologically diverse languages
| # | LLM | Retriever | Type | Judge | EM Overall ▾ | Recall Overall | Search Turns Avg. |
|---|---|---|---|---|---|---|---|
| 1 | gpt-oss-20b Baseline | Qwen3-8B |
Agentic | Opus 5 | 39.47 | 54.34 | 27.79 |
| 2 | Tongyi-DR-30B Baseline | Qwen3-8B |
Agentic | Opus 5 | 39.07 | 47.57 | 26.04 |
| 3 | DR-Tulu-8B Baseline | Qwen3-8B |
Agentic | Opus 5 | 13.33 | 24.19 | 4.77 |
| Language | Tier | gpt-oss-20b | Tongyi-DR-30B | DR-Tulu-8B | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| EM (%) | Recall (%) | Search Turns | EM (%) | Recall (%) | Search Turns | EM (%) | Recall (%) | Search Turns | ||
English en | High | 48.00 | 62.83 | 28.46 | 50.00 | 64.36 | 30.60 | 32.00 | 43.63 | 5.00 |
Chinese zh | High | 44.00 | 61.85 | 27.26 | 38.00 | 40.71 | 23.28 | 12.00 | 21.51 | 4.72 |
French fr | High | 44.00 | 60.08 | 28.16 | 42.00 | 48.94 | 28.18 | 18.00 | 31.94 | 4.62 |
Russian ru | High | 44.00 | 60.29 | 28.98 | 36.00 | 51.93 | 26.66 | 18.00 | 30.53 | 4.84 |
Spanish es | High | 40.00 | 48.37 | 27.70 | 46.00 | 50.81 | 28.06 | 16.00 | 27.33 | 4.62 |
German de | High | 44.00 | 55.68 | 26.00 | 36.00 | 46.69 | 26.12 | 10.00 | 23.97 | 4.86 |
Arabic ar | Mid | 42.00 | 55.57 | 28.12 | 52.00 | 52.92 | 26.08 | 12.00 | 26.56 | 4.56 |
Bengali bn | Mid | 40.00 | 55.12 | 27.92 | 44.00 | 54.72 | 26.72 | 16.00 | 21.82 | 4.78 |
Finnish fi | Mid | 34.00 | 51.77 | 28.88 | 34.00 | 50.16 | 29.92 | 14.00 | 23.22 | 5.06 |
Hindi hi | Mid | 40.00 | 56.68 | 27.72 | 48.00 | 51.03 | 25.12 | 16.00 | 26.20 | 4.82 |
Thai th | Mid | 32.00 | 50.97 | 26.66 | 30.00 | 40.56 | 19.88 | 8.00 | 23.85 | 4.88 |
Urdu ur | Mid | 38.00 | 51.63 | 26.88 | 36.00 | 51.56 | 27.48 | 12.00 | 20.81 | 4.60 |
Tamil ta | Mid | 42.00 | 56.00 | 29.70 | 46.00 | 48.94 | 25.06 | 10.00 | 21.79 | 4.52 |
Swahili sw | Low | 28.00 | 44.96 | 27.16 | 24.00 | 29.84 | 25.28 | 2.00 | 7.29 | 4.92 |
Welsh cy | Low | 32.00 | 43.31 | 27.26 | 24.00 | 30.37 | 22.10 | 4.00 | 12.39 | 4.78 |
EM and recall are reported as percentages over the held-out test queries; search turns are average search calls per query. Each overall value is the simple average of its per-language values. Teams that submit runs but do not submit a working note may be excluded from the final leaderboard. Where a team submits multiple runs for a language, only the best-scoring eligible run is ranked.
JSONL output format shown below. Example submissions available on GitHub → BrowseComp-Plus.
{
"query_id": "zh-798",
"language": "chinese",
"retriever": "Qwen/Qwen3-Embedding-8B",
"llm": "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B",
"tool_call_counts": { "search": 24 },
"retrieved_docids": [
["81120", "10986", "11064", "83777", "9823"],
...
["81120", ...] // kth search round
],
"result": [
{
"type": "reasoning",
"tool_name": null,
"arguments": null,
"output": "Let me search for information about..."
},
{
"type": "tool_call",
"tool_name": "search",
"arguments": "query: \"As of December 2023\" coordinator research group founded in 2009",
"output": "[{\"docid\": \"81120\", \"score\": 8.43, \"snippet\": \"...\"}]"
},
{
"type": "output_text",
"tool_name": null,
"arguments": null,
"output": "Explanation: I attempted to solve this ... Exact Answer: compassionate pugilist"
}
]
}
No more than 3 output files per language per team. There is no restriction on the language used by the LLM during reasoning; however, direct translation of the original query to English as the only retrieval strategy is discouraged.
Validate offline, then upload through the submission portal.
Every file is one language: one JSON object per line covering all 50 official query ids of that language, with the same language, llm and retriever in every record. Check files before uploading with mast-validate; the portal runs exactly the same checks.
pip install git+https://github.com/mast-benchmark/mast-validate
mast-validate runs/hi.jsonl --track indic
mast-validate runs.zip --track multilingual --json report.json
Exit code 0 is clean, 1 warnings only, 2 errors. Errors block a file; warnings do not, but each one costs score.
Submit at the MAST 2026 submission portal with your registered team name, the track, and an email address from your registration. Upload one .jsonl (or .jsonl.gz), or one .zip / .tar.gz holding one file per language, named however you like.
One button: the page validates the upload and records every file that passes, then shows one row per language with what happened and a receipt. Files with errors are not recorded; fix them and submit again, and unchanged languages are skipped automatically.
Three slots per team, track and language. A new run takes the next free slot; when all three are filled, tick replace the oldest run to overwrite the oldest. Every recorded run gets a receipt id and a sha256; keep the receipt, it is your proof of submission and exactly what gets evaluated.
Query ids may be the dataset form ("zh-798") or a bare number such as 798; both are accepted. Deadline: September 15, 2026, 23:59 AoE.
System description papers for FIRE 2026 proceedings
Each participating team must submit one working note per track describing their system and approach. Working notes are submitted centrally via the FIRE 2026 submission system.
Working notes must follow the CEUR-WS single-column format (ceurart style, as used by the CEUR Workshop Proceedings). Minimum 6 pages, maximum 10 pages, including references. Use descriptive titles — do not include team or track names in the title.
⚠️ Author names should not include prefixes (Dr., Prof., etc.). At least one author per working note must be registered for FIRE 2026. Teams that submit runs but do not submit a working note may be excluded from the final leaderboard.
MAST @ FIRE 2026 Tentative Timeline (subject to change)
Three steps to join MAST
Fill in the registration form with your team name, member names, affiliations, and the language(s) you plan to participate in. Teams are limited to 4 members.
Register →Access the released test queries and the English document corpus through the MAST Hugging Face organization. Baseline retrieval runs will be provided.
Multilingual Queries → Indic Queries → Documents →Validate your JSONL files with mast-validate, then upload up to 3 runs per language through the submission portal before September 15, 2026. Submit a working note describing your system and approach (6–10 pages including references, CEUR-WS format).
Registration is free and open to academic groups, industry labs, and individual researchers. After registering, you will be added to the track mailing list and Slack workspace.
Registration closes September 15th, 2026.

Lead of BEIR and MIRACL. Co-organizer of TREC-RAG 2024 & 2025.

Lead of Mr. TyDi and MIRACL. Co-organizer of CIRAL @ FIRE 2023.
Building on a strong foundation