Domain: Manufacturing & Industrial Robotics (pick-and-place, assembly, logistics) · related: Vision-Language Models, Embodied AI, Industrial AI
Industrial robots excel at repeating fixed programs in rigid environments but struggle when tasks change frequently. In small-batch Swiss manufacturing, operators constantly adapt to new orders. Natural language is the most intuitive interface for this variability — but cloud AI is a non-starter: under strict IP confidentiality and the revised Data Protection Act (revDSG), Swiss companies cannot stream production imagery to foreign servers. The intelligence must live on the factory floor. This challenge tests whether Apertus 1.5 — Switzerland's open model — can serve as the cognitive brain for sovereign, edge-computed industrial robotics. ZHAW opens its physical hardware: a UR5 workcell with RGB-D and LiDAR sensing, a calibrated camera, and a proven grasping pipeline (six years of research). Your task is to build the missing cognitive layer that turns plain-English instructions and raw camera feeds into safe, executable robot action plans.
Programming a robot arm is a specialist job, and making it natural requires grounding language in reality — but LLMs hallucinate physics. Teams use an open VLM to turn a camera feed into a scene graph, then prompt Apertus to translate an instruction (e.g. "Clear the table but leave the blue tray") into a safe, ordered list of actions. Three specific failures matter:
Build a zero-shot pipeline from image + text to robot action, using an open vision model together with Apertus. No training needed. Input: one workcell photo + one English instruction. Output: a JSON object with the pixel position of the target object, a short ordered action list, or a question / refusal. The pixel position is key — the challenge provides camera parameters and hand-eye calibration, so a correct pixel converts directly into a robot coordinate for execution on the physical arm after the event. Recommended (not required): open VLM → structured scene graph → prompt Apertus (tool-calling / constrained decoding) → ordered action list. Teams may also try Apertus alone via its image input — comparing the two is one of the most useful results. No robotics knowledge and no robot are required.
Raw sensor snapshots from a physical, calibrated robot cell (teams work zero-shot; no training):