Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation

X. He, T.-E. Kim, M. Fröebe, J. Arguello, B. Mitra, F. Diaz
SIGIR, 2026
Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM–based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English.

bibtex

Copied!
@inproceedings{he:multilingual-tot-synth, year = {2026}, title = {Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation}, booktitle = {Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval}, author = {Xuhong He and To Eun Kim and Fernando Diaz and Maik Fr{\"o}ebe and Jaime Arguello and Bhaskar Mitra} }