IRLBench: A multi-modal, culturally grounded, parallel Irish-English benchmark for open-ended LLM reasoning evaluation

dc.contributor.authorTran, Khanh Tung
dc.contributor.authorNguyen, Duc Hai
dc.contributor.authorO'Sullivan, Barry
dc.contributor.authorNguyen, Hoang D.
dc.contributor.funderTaighde Éireann - Research Ireland
dc.contributor.funderEuropean Regional Development Fund
dc.date.accessioned2026-05-22T11:50:01Z
dc.date.available2026-05-22T11:50:01Z
dc.date.issued2026-04-20
dc.description.abstractRecent advances in Large Language Models (LLMs) have demonstrated promising capabilities, yet their performance in multilingual and low-resource settings remains modest. Existing benchmarks often exhibit cultural bias, restrict evaluation to text-only, rely on multiple-choice formats, and, more importantly, are ineffectual for extremely low-resource languages. To address these gaps, we introduce IRLBench, presented in parallel English and Irish, which is considered definitely endangered by UNESCO. Our benchmark consists of 12 representative subjects developed from the 2024 Irish Leaving Certificate exam, enabling fine-grained analysis of model capabilities across domains. By framing the task as long-form generation and leveraging the official marking scheme, it supports not only a comprehensive evaluation of correctness but also language fidelity. Our extensive experiments of leading closed-source and open-source LLMs reveal a persistent performance gap between English and Irish, in which models produce valid Irish responses less than 80% of the time, and answer correctly 55.8% of the time compared to 76.2% in English for the best-performing model. With Irish as the case study, our work exposes systemic weaknesses in today's multilingual LLMs and provides a rigorous benchmark for evaluating true multilingual capabilities. We release IRLBench and an accompanying evaluation codebase to enable future research on robust, culturally aware multilingual AI development.en
dc.description.sponsorshipThis publication has emanated from research supported in part by grants from Research Ireland under Grant [12-RC-2289-P2] and [18/CRT/6223] which is co-funded under the European Regional Development Fund.
dc.description.statusPeer reviewed
dc.description.versionPublished Version
dc.format.extent12
dc.format.extent3136678
dc.format.mimetypeapplication/pdf
dc.identifier.authororcidTran, Khanh Tung
dc.identifier.authororcidNguyen, Duc Hai
dc.identifier.authororcidO'Sullivan, Barry§0000-0002-0090-2085
dc.identifier.authororcidNguyen, Hoang D.§0000-0003-2541-3269
dc.identifier.citationTran, K. T., Nguyen, D. H., O'Sullivan, B. and Nguyen, H. D. (2026) 'IRLBench: A multi-modal, culturally grounded, parallel Irish-English benchmark for open-ended LLM reasoning evaluation', Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD (2026) Jeju Island, Republic of Korea, 9 August 2026, pp. 2794-2805. https://doi.org/10.1145/3770854.3785693
dc.identifier.doi10.1145/3770854.3785693
dc.identifier.endpage2805
dc.identifier.isbn9798400722585
dc.identifier.issn2154-817X
dc.identifier.journaltitleKDD 2026 - Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1
dc.identifier.journaltitle32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026
dc.identifier.otherORCID: /0000-0003-2541-3269/work/215517875
dc.identifier.startpage2794
dc.identifier.urihttps://hdl.handle.net/10468/18832
dc.identifier.urlhttps://www.scopus.com/pages/publications/105038096486
dc.language.isoen
dc.publisherAssociation for Computing Machinery
dc.relation.ispartofseriesProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
dc.rights© 2026, Owner/Author. For the purpose of Open Access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission.
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subjectArtificial Intelligence and Data Analytics
dc.subjectBenchmarking
dc.subjectExtremely low-resource
dc.subjectLarge language model
dc.subjectLarge vision-language model
dc.subject[ComputerScience]
dc.subject[Insight Centre for Data Analytics]
dc.subjectSoftware
dc.subjectInformation Systems
dc.titleIRLBench: A multi-modal, culturally grounded, parallel Irish-English benchmark for open-ended LLM reasoning evaluationen
dc.typeConference item
Files
Original bundle
Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
3770854.3785693.pdf
Size:
2.99 MB
Format:
Adobe Portable Document Format