Century Automation← All news

AI & Automation Briefing - October 1, 2026

OpenAI Launches Decisions API to Bring Fast, Cheap Structured Decisions to Agent Workflows

At its Dev Day event, OpenAI announced a Decisions API that functions similarly to Jev, a model released by TypeSafe AI earlier in September. Jev is built on an LLM but designed specifically for software automation. It accepts a predefined set of choices and returns them as probabilities at high speed and low cost, making it a practical alternative to full LLM calls for routing and classification tasks. OpenAI's version lets its Luna model select from a fixed set of options, including image categories or agent behaviors, with speed and cost advantages baked in. Sam Altman described it as a way to keep image understanding, language support, and safety protections intact while dramatically narrowing the model's decision scope. TypeSafe CEO Diogo Almeida, a former OpenAI engineer, noted on X that the move signals "building in a System One compatible way is the future," using his company's term for fast, intuitive processing as opposed to slower deliberate reasoning. The practical implication for agent builders is significant. Most LLM calls in multi-agent pipelines are overbuilt for simple routing decisions. A dedicated decision model handles those branches faster and cheaper, which changes how orchestration layers like n8n workflows and multi-agent systems should be architected. OpenAI released the Decisions API as a limited preview, and calibration quality across competing products remains an open question. Almeida argues TypeSafe's edge is synthetic data quality, not just speed. Several other startups are shipping similar models, and OpenAI is unlikely to be the last large lab to enter this category.

Source

Restate Raises $20M Series A to Solve Durable Execution for AI Agent Workflows

Berlin-based Restate has closed a $20 million Series A led by Singular, with Redpoint Ventures and Capital One Ventures participating. The startup builds an execution engine that keeps multi-step workflows running reliably through crashes, network failures, and mid-process errors, tracking each step so outcomes stay consistent even when something goes wrong. That capability maps directly onto the problems AI agent workflows create, since agents run longer and follow less predictable paths than traditional software. Restate was not originally built for agents, but the fit became obvious as agentic use cases scaled. Unlike competitors, Restate built its own storage, replication, and redundancy layers rather than sitting on top of an external database, which the company says makes it faster and cheaper to run than heavier durable-execution systems. Customers include Replit and several Fortune 500 financial firms. The primary competitor is Temporal, which raised $550 million at a $12.55 billion valuation earlier this month. Restate was co-founded by Stephan Ewen, a co-creator of Apache Flink, along with former Data Artisans and Ververica colleagues Igal Shilman and Till Rohrmann.

Source

Galahad Caches LLM Attention State to Disk, Turning Repeated Document Processing Into a One-Time Cost

Researchers found that across seven real-world serving datasets, 98.7% of prompt tokens were text the model had already processed in a prior request. Because standard inference is stateless, models recompute the full attention state every time the same document appears in a new prompt. Galahad, a memory layer built for vLLM, SGLang, and llama.cpp, addresses this by saving the key-value cache for any block of text to disk and reloading it on the next matching request rather than recomputing it. The system has two components: Taliesin handles KV cache persistence and restoration, and Blaise handles document storage and retrieval, passing only the relevant section to the model per question. On a 97,000-token recall benchmark using Gemma 4 31B, the combined system answered 100 of 100 questions at 0.59 to 0.64 seconds and 200 to 213 joules per question, compared to 10 of 100 for the same model without Galahad, which could only fit the last 12,000 tokens in context. Restored state is bit-identical to what the model would have computed directly. The upfront cost of storing a corpus is roughly 100 seconds and 28 kilowatt-joules, which the system recoups after approximately 13 questions. For workflows that repeatedly send the same documents through LLM nodes, this architecture directly reduces both latency and compute cost.

Source

Sources