Wire Observer.
Technology

Engineers Test Three Leading AIs in a Real‑World Drive, Only One Reaches the Fast‑Food Destination

Engineers Test Three Leading AIs in a Real‑World Drive, Only One Reaches the Fast‑Food Destination

In a hands‑on experiment that blended cutting‑edge language models with everyday automotive technology, three engineers tasked three prominent AI systems—OpenAI's GPT, Anthropic's Claude, and xAI's Grok—with the challenge of piloting a Toyota Corolla to a nearby In‑N‑Out restaurant. The test aimed to assess whether conversational models, typically confined to text, could translate their reasoning into safe, real‑time vehicle control.

The team equipped the Corolla with a standard suite of sensors, including cameras and lidar, and connected each AI to a custom interface that relayed perception data and accepted steering, throttle, and braking commands. While the models were not originally designed for motor control, the engineers provided them with a brief on‑board instruction set that described how to interpret sensor inputs and issue driving actions.

During the trial, GPT managed to navigate the streets but struggled with dynamic obstacles, hesitating at intersections and occasionally issuing contradictory commands that required manual override. Claude demonstrated a more cautious approach, maintaining lane discipline but failing to progress toward the destination, effectively getting stuck in a loop of indecision at traffic signals. In contrast, Grok succeeded in completing the entire route without human intervention, smoothly handling turns, merges, and the final stop at the fast‑food outlet.

Analysts note that the differing outcomes highlight how each model's training emphasis influences practical performance. GPT's strength lies in expansive language generation, which does not automatically translate to split‑second decision‑making required for driving. Claude's design prioritizes safety and interpretability, leading to overly conservative behavior in a fluid traffic environment. Grok, built with a more integrated approach to multimodal tasks, appears better suited to converting textual reasoning into actionable motor commands.

The experiment underscores both the promise and the current limitations of repurposing large language models for embodied AI tasks. While Grok's success suggests a pathway toward more capable autonomous systems, the need for robust sensor integration, real‑time processing, and rigorous safety validation remains critical. The engineers plan to refine the interface, expand testing to varied road conditions, and explore how hybrid models might combine the linguistic depth of GPT with the operational reliability demonstrated by Grok.

Source: wired
Aarav Mehta — Technology desk.

Comments (0)

Be the first to comment.

Join the discussion

Protected by reCAPTCHA v3

Related