Extracting Parliamentary Affairs from PDFs into One Common Structure

🎯 Challenge Overview

Domains: Public Sector / GovTech / Democracy / Politics

→ Context & Background

Switzerland has parliaments at three levels — the Federal Assembly, 26 cantonal parliaments, and several hundred municipal/city parliaments — operating across different languages. Most proceedings involve a common object: a parliamentary affair (motion, postulate, interpellation, written question or initiative), submitted by members, answered by the executive, and usually closed by a decision. OpenParlData (openparldata.ch) scrapes 90 Swiss parliaments daily and publishes ~300,000 parliamentary affairs as open data, via API (api.openparldata.ch) and bulk exports (files.openparldata.ch/exports/). Much valuable information is buried in parliamentary PDFs and can't be reliably extracted into structured form — the question asked, the government's answer and reasoning, the decision, the people involved. Documents vary hugely: born-digital, scanned, some with handwritten stamps/signatures, tables and annexes. Generic document AI (e.g. Docling) gives a clean but domain-agnostic page representation; ad-hoc LLM prompting produces a different JSON shape for every body. OpenParlData already runs the surrounding infrastructure (a prompt-driven extraction workbench with side-by-side PDF / Text / JSON / Docling views and a manual review step), so what this challenge produces can go into real operation, not just a demo.

→ Problem Description

There is no shared, machine-readable format in which a parliamentary affair can be expressed so two parliaments become comparable — and no reliable way to produce one from existing PDFs. Three gaps:

  1. Format gap. Docling/plain text are too generic (document structure, not affair semantics); Akoma Ntoso is rich but heavyweight and rooted in legislative drafting; TEI/Parla-CLARIN targets speech/debate transcripts. What's missing is a middle layer (like JATS for scientific papers) that preserves document structure and captures affair semantics (type, submitter & co-signatories, addressed body, question/answer pairings, dates, procedural history, decisions, links to related affairs).
  2. Extraction gap. Turning heterogeneous, multilingual, partly-scanned PDFs into that format at 300,000-document scale needs an LLM — and an unconstrained, untraceable LLM invents values.
  3. Verification gap. No gold standard and no metric, so nobody can say whether a pipeline is good enough to run unattended.

→ Primary Objective

Design and prototype a structured representation model for parliamentary affairs, plus a working Apertus-driven pipeline that produces it from real PDFs. Three deliverables:

🔧 Resources, Tools & Support

→ Datasets

Provided by OpenParlData: