Century Automation← All news

AI & Automation Briefing - June 28, 2026

Ford Rehires Veteran Engineers After AI Quality Tools Fall Short

Ford has brought back roughly 350 experienced engineers over the past three years to retrain younger staff and reconfigure AI inspection tools that failed to meet quality standards. Many of the rehired engineers are former Ford employees or came from suppliers. The move appears to have paid off: Ford ranked as the top mainstream brand in the 2026 JD Power Initial Quality Survey. The case is a direct example of over-automation creating a regression that required human expertise to correct, and a concrete illustration of why human-in-the-loop design matters before full production deployment.

Source

U.S. Lifts Block on Anthropic's Claude Mythos 5, Releases It to Over 100 Trusted U.S. Entities

The U.S. Commerce Department reversed its export controls on Anthropic's Claude Mythos 5 model on June 27, permitting access for more than 100 U.S. companies and government agencies. Commerce Secretary Howard Lutnick notified Anthropic in a letter that the department had determined sufficient safeguards were in place, following two weeks of daily negotiations after the original block was imposed over concerns the model could be jailbroken for malicious use. The arrangement eliminates the license requirement for approved entities and their foreign national employees. A companion model, Fable 5, remains under restriction, though people close to the talks say a release is in progress on an unclear timeline. The move signals the early formation of a federal framework giving the government direct oversight over frontier AI model releases. OpenAI released GPT-5.6 to a short list of government-approved partners on the same day.

Source

New Benchmark Exposes Where AI Agents Break Down Outside Familiar Tasks

A research paper introducing the Gauntlet benchmark tests AI agents on web-based tasks requiring temporal perception, graphical interpretation, and 3D spatial reasoning. Agents scored significantly below human performance across all three categories, revealing that current systems struggle to generalize beyond environments they were trained on. For automation builders, this data points to specific failure modes worth accounting for in workflow design, including where to build fallback logic and human checkpoints. The benchmark also provides citable evidence for setting realistic expectations with clients about where agentic systems will reliably underperform.

Source

Sources