10 articles tagged with “benchmarks”.
AI NewsAn AWS post introduces Self-Distilled Reasoning (SDR), an approach for adding reasoning traces to supervised fine-tuning datasets that lack them, validated across three benchmarks with Amazon Nova.
AI NewsOpenAI CFO Sarah Friar presents a practical scorecard for evaluating AI ROI, based on useful work, cost per successful task, dependability, and return on compute.
AI NewsA survey of 101 enterprises finds AI agents often produce confident but wrong answers traced to missing or inconsistent business context, exposing a trust gap in enterprise RAG infrastructure.
AI NewsNVIDIA's Nemotron 3 Embed model has taken the top overall spot on the RTEB benchmark, positioned around advancing agentic retrieval.
AI NewsOpenAI created GPT-Red, an LLM designed to act as an attacker that helps train its models against cyberattacks. The company says GPT-5.6 is its most robust release yet after such training.
AI NewsAWS details how Thrad.ai deployed a multi-agent system using Strands Agents and Amazon Bedrock AgentCore to automate prospect discovery and personalized email generation, with benchmarks on two orchestration patterns.
AI NewsAWS has launched a UI in Amazon SageMaker AI Studio that guides users to optimized generative AI inference configurations without deep infrastructure expertise.
GuidesLetting a model 'think' before answering measurably improves hard reasoning. Here is how chain-of-thought works, how it grew into dedicated reasoning models, and when the extra cost pays off.
GuidesEvery word an LLM generates has a cost in compute, memory, and time. Here is what actually happens during inference — and why it explains latency, throughput, and per-token pricing.
GuidesA demo that works on five hand-picked prompts is not a working AI product. Here is how to measure LLM quality honestly — the metrics, the judge models, and the datasets — so you catch failures before your users do.