This track 1 is focused on improving the multimodal Apertus v1.5 8B and 70B models. There are two different paths. Path B focuses on localising Apertus to the languages, dialects, knowledge and values of Switzerland. Submissions can focus either on the 8B or 70B model.
The submission must comply with the HackApertus T&C. For this challenge it is essential that you pay particular attention to the conditions for responsibly sourced datasets.
Localised alignment to the languages, dialects, knowledge and values of Switzerland
👉🏼You can pick one localisation dimension for your submission.
This challenge accepts submissions that evaluate any of the following dimensions of localising Apertus to specific Swiss contexts: input adaptation, domain-specific core task intelligence, and behavioural alignment. These dimensions are motivated as follows.
Many speech-native multimodal LLMs struggle with phonetic drift, code-switching, local accents, and unwritten or low-resource dialects. Submissions should measure transcription accuracy and dialect comprehension to evaluate Apertus’s ability to transcribe Swiss voice inputs into written text representations without losing local context. Note that audio input is an experimental feature in Apertus 1.5.
Dialects that belong to the following ISO-639-3 dialect groups are permitted: gsw, wae, roh, lmo, frp. Dialects should be reported with their glottocodes.
General base models usually struggle with domain-specific knowledge and niche terminology. Submissions should measure factuality and hallucination rates to evaluate whether Apertus possesses factually accurate, domain-tailored knowledge to execute specialised reasoning for localised Swiss contexts (e.g., legal statutes, local medical guidelines, local government administration processes).
Model performance is not only about dialects and facts, but also about cultural nuances, tone, safety boundaries, and local ethics. For every item in the dataset, the "correct" behaviour must trace to a source of normative authority. This could be a codified law or regulation, a documented professional or institutional standard, a measured majority preference (primary data collection or citation), a convention documented in a style guide or reference work. An item whose only warrant is the author's intuition does not present an authoritative source.
Submissions should measure how frequently responses meet end user expectations, to evaluate which regional norms and local regulatory standards Apertus respects by default.
Each localisation approach calls for a different evaluation dataset type and has its own expectations for test items and evaluation metrics. The table below lists minimum dataset expectations for each localisation approach and example evaluation metrics. We expect the winning submissions to present datasets and evaluations that are significantly deeper than the minimum requirements.
You should describe and motivate your dataset and evaluation design choices in detail in the submission report.
| Localisation approach | Minimum evaluation dataset expectations | Example evaluation metrics |
|---|---|---|
| Input Adaptation | 100 curated audio clips, paired with text transcriptions and translations to a standard written language (e.g. German, French, Italian, Romansh) | Transcription: mean character error rate and word error rate across all audio clips. |
| Translation: BLEU, chrF, COMET. | ||
| Core Task Intelligence | 300 single-turn Q&A pairs with answer variations and source citations | SimpleQA-style three-way grading (correct/incorrect/not attempted) against ground-truth references using LLM-as-a-judge (with a random sample of 20% of judgements evaluated by a human and reported judge–human agreement). |
| Behavioural Alignment | 200 scenario prompts with explicit criteria defining acceptable vs. unacceptable answers, and a matched pair of chosen vs. rejected outputs; must be anchored in a specific use case and user community; expected behaviours should be labelled as contested vs. settled and treated accordingly. |
Forced-choice preference accuracy, rubric-scored free generation with LLM-as-a-judge scoring against stated acceptance criteria (with a random sample of 20% of judgements evaluated by a human and reported judge–human agreement). |