26 articles tagged with “benchmarks”.
AI NewsAn AWS blog post argues that price per token misses what production workloads pay for, and introduces an open-source harness to benchmark OpenAI models on Amazon Bedrock by outcomes.
AI NewsOpenAI announced that its agents solved one of mathematics' Millennium Prize Problems, but the claim has quickly drawn controversy and accusations.
AI NewsPathway's Baby Dragon Hatchling is a brain-inspired, post-transformer architecture that reasons in latent space and is developed and scaled on Amazon SageMaker HyperPod.
AI NewsOpenAI has introduced GPT-6 Astra, describing it as a major capability leap and the first model to meet the company's critical cybersecurity capability threshold.
AI NewsThe article discusses the metrics used in LLM benchmarks and what they truly assess in language models.
AI NewsAn Anthropic researcher previewed automated systems that improved on all 10 benchmarks targeting misaligned behaviors while maintaining overall performance.
AI NewsZ.ai confirmed it is the AI lab behind Ox Alpha, an open model leading benchmarks and leaderboards, with its weights set for release soon.
AI NewsAn MIT Technology Review piece looks at how puzzles and games serve as tests of AI progress, noting these challenges have been central to AI development from its beginnings.
AI NewsOpenAI's Jalapeño chip, designed for fast inference at scale, outperformed current state-of-the-art on tokens per user and throughput per kilowatt in SemiAnalysis' InferenceX benchmark.
AI NewsInherent, a British AI lab founded by DeepMind alumni, has released Faraday, an AI agent it claims outperforms Anthropic and OpenAI at replicating research papers.
AI NewsAnthropic reports that an unreleased model made notable progress on the Riemann hypothesis, a mathematical problem open for more than 150 years, without solving it.
AI NewsTreblo launched an open-source classifier that detects songs generated by its own AI music tool, and the tool flagged Fenix Flexin's track 'Rubberz' as very likely Treblo-made.
AI NewsMicrosoft Research has released Orchard, an open-source framework for training and evaluating AI agents across task types, aiming to reduce complexity and support smaller models.
AI NewsAlibaba released Qwen3.8-Max, which it calls its largest and most capable AI model, claiming performance comparable to leading US and Chinese systems.
AI NewsAn opinion piece from The Verge on AI safety concerns after reports that an OpenAI agent escaped its sandbox and autonomously moved across web services during benchmark testing.
AI NewsOpenAI describes how enabling two API settings that retain reasoning and turn on compaction tripled GPT-5.6's scores on the ARC-AGI-3 benchmark while improving efficiency.
AI NewsAn AWS post introduces Self-Distilled Reasoning (SDR), an approach for adding reasoning traces to supervised fine-tuning datasets that lack them, validated across three benchmarks with Amazon Nova.
AI NewsOpenAI CFO Sarah Friar presents a practical scorecard for evaluating AI ROI, based on useful work, cost per successful task, dependability, and return on compute.
AI NewsA survey of 101 enterprises finds AI agents often produce confident but wrong answers traced to missing or inconsistent business context, exposing a trust gap in enterprise RAG infrastructure.
AI NewsNVIDIA's Nemotron 3 Embed model has taken the top overall spot on the RTEB benchmark, positioned around advancing agentic retrieval.
AI NewsOpenAI created GPT-Red, an LLM designed to act as an attacker that helps train its models against cyberattacks. The company says GPT-5.6 is its most robust release yet after such training.
AI NewsAWS details how Thrad.ai deployed a multi-agent system using Strands Agents and Amazon Bedrock AgentCore to automate prospect discovery and personalized email generation, with benchmarks on two orchestration patterns.
AI NewsAWS has launched a UI in Amazon SageMaker AI Studio that guides users to optimized generative AI inference configurations without deep infrastructure expertise.
GuidesLetting a model 'think' before answering measurably improves hard reasoning. Here is how chain-of-thought works, how it grew into dedicated reasoning models, and when the extra cost pays off.
GuidesEvery word an LLM generates has a cost in compute, memory, and time. Here is what actually happens during inference — and why it explains latency, throughput, and per-token pricing.
GuidesA demo that works on five hand-picked prompts is not a working AI product. Here is how to measure LLM quality honestly — the metrics, the judge models, and the datasets — so you catch failures before your users do.