See It, Say It, Pick It — Vision-Language Grounding for a Real Robot Arm

🎯 Challenge Overview

Domain: Manufacturing & Industrial Robotics (pick-and-place, assembly, logistics) · related: Vision-Language Models, Embodied AI, Industrial AI

→ Context & Background

Industrial robots excel at repeating fixed programs in rigid environments but struggle when tasks change frequently. In small-batch Swiss manufacturing, operators constantly adapt to new orders. Natural language is the most intuitive interface for this variability — but cloud AI is a non-starter: under strict IP confidentiality and the revised Data Protection Act (revDSG), Swiss companies cannot stream production imagery to foreign servers. The intelligence must live on the factory floor. This challenge tests whether Apertus 1.5 — Switzerland's open model — can serve as the cognitive brain for sovereign, edge-computed industrial robotics. ZHAW opens its physical hardware: a UR5 workcell with RGB-D and LiDAR sensing, a calibrated camera, and a proven grasping pipeline (six years of research). Your task is to build the missing cognitive layer that turns plain-English instructions and raw camera feeds into safe, executable robot action plans.

→ Problem Description

Programming a robot arm is a specialist job, and making it natural requires grounding language in reality — but LLMs hallucinate physics. Teams use an open VLM to turn a camera feed into a scene graph, then prompt Apertus to translate an instruction (e.g. "Clear the table but leave the blue tray") into a safe, ordered list of actions. Three specific failures matter:

  1. Visual grounding — connect words to one object at one position (VLMs name what's in an image but are weak on where).
  2. Order & physics — one sentence can mean several actions in a required order; the gripper holds one object at a time.
  3. Ambiguity & refusal — some instructions are under-specified, name absent objects, or are physically impossible. The correct answer is then a question or a refusal, not a plan. In ZHAW's pilot, Apertus produced a confident plan for every case — including for an object that wasn't on the table. For a machine that moves, that's the failure that matters most.

→ Primary Objective

Build a zero-shot pipeline from image + text to robot action, using an open vision model together with Apertus. No training needed. Input: one workcell photo + one English instruction. Output: a JSON object with the pixel position of the target object, a short ordered action list, or a question / refusal. The pixel position is key — the challenge provides camera parameters and hand-eye calibration, so a correct pixel converts directly into a robot coordinate for execution on the physical arm after the event. Recommended (not required): open VLM → structured scene graph → prompt Apertus (tool-calling / constrained decoding) → ordered action list. Teams may also try Apertus alone via its image input — comparing the two is one of the most useful results. No robotics knowledge and no robot are required.

🔧 Resources, Tools & Support

→ Datasets

Raw sensor snapshots from a physical, calibrated robot cell (teams work zero-shot; no training):