Up to 3.2x Faster Inference with LFM2.5-DSpark
The introduction of LFM2.5-DSpark enhances inference speed by up to 3.2 times, benefiting developers in AI.
34 articles tagged with “inference”.
The introduction of LFM2.5-DSpark enhances inference speed by up to 3.2 times, benefiting developers in AI.
This article discusses the memory demands of AI agents, emphasizing the importance of optimizing memory usage for enhanced performance.
NVIDIA Nemotron 3.5 Lightning, an open 30B Mixture-of-Experts model with 3B active parameters, is now available in Amazon SageMaker JumpStart for high-volume agentic workloads.
AWS describes building multi-agent workflows using OpenAI-compatible SageMaker AI endpoints and Bedrock AgentCore runtime, with each agent using a model suited to its task.
Writer introduced a new AI model and an upgraded harness designed to contain token costs, built as a post-training variation on Z.ai's open source GLM-5.2 model.
OpenAI is previewing Ultrafast, an API service tier that runs GPT-5.6 Sol up to 14x faster, reaching up to 750 output tokens per second using Cerebras hardware.
An AWS blog post details building a tiered KV cache on Amazon SageMaker HyperPod with Curvine, extending the cache into a shared NVMe pool so replicas reuse cache on cost-efficient instances.
NVIDIA has added Nemotron 3.5 Lightning to its Nemotron 3 model family, alongside NeMo Switchyard, aimed at efficient, long-running agentic AI workloads with open models.
Anthropic, the maker of Claude, is building a team to design its own custom AI chips, aiming to co-design hardware and models for improved speed and efficiency.
NVIDIA highlights how growing AI demands push datasets and context windows past system memory, requiring efficient, secure storage architectures rather than added capacity alone.
A Hugging Face Blog post introduces LFM2.5-2.6B, presented as a way to deploy local AI agents across environments.
OpenAI details GPT-Live, a realtime voice AI system built over six months that uses a turnless speech model and low-latency architecture for continuous, more natural voice interaction.
OpenAI describes a full-stack strategy intended to make advanced AI more capable, more affordable, and more broadly useful.
OpenAI GPT-5.6 Sol, Terra, and Luna are generally available on Amazon Bedrock, now with explicit prompt caching that lets users choose which prompt parts to cache and reuse.
OpenAI announced lower GPT-5.6 pricing for its Luna and Terra offerings, positioning more efficient models to support enterprise AI workflows at scale.
OpenAI describes how enabling two API settings that retain reasoning and turn on compaction tripled GPT-5.6's scores on the ARC-AGI-3 benchmark while improving efficiency.
The OlmoEarth platform is designed for geospatial inference, enabling advanced data analysis across the globe.
Satya Nadella cautions that businesses depending on one AI for everything risk trouble unless they build their own models or use AI gateways to separate prompts from the underlying model.
AWS introduces Claude Opus 5, Anthropic's most capable Opus model, with practical guidance for engineers building agentic systems and production inference workloads on Amazon Bedrock.
Three OpenAI GPT-5.6 models—Sol, Terra, and Luna—are now generally available on Amazon Bedrock, with support for the Responses API, prompt caching, and the Codex coding agent.
Google has released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, while notably not releasing a Gemini 3.5 Pro model.
NVIDIA's Vera Rubin is moving into production, with NVL72 racks deployed at major cloud partners and backed by a supply chain spanning over 350 factory sites in 30 countries.
NVIDIA's Vera Rubin is positioned to reduce cost per token for post-training workloads through codesign, maximizing intelligence per dollar for agentic AI.
A $400 million chip-backed loan highlights a move by early GPU financiers toward inference chips as part of the next wave of AI infrastructure financing.
NVIDIA unveiled the T3000 and T2000 modules based on its Thor architecture, aimed at bringing compact, power-efficient AI computing to mass-market robotics and edge AI.
NVIDIA frames performance per watt as the defining efficiency metric for AI infrastructure, arguing power is the main constraint on how many tokens an AI factory can generate.
OpenAI's GPT-5.6 Sol, Terra, and Luna models are now generally available on Amazon Bedrock, accessible through its inference engine.
AWS has launched a UI in Amazon SageMaker AI Studio that guides users to optimized generative AI inference configurations without deep infrastructure expertise.
Every word an LLM generates has a cost in compute, memory, and time. Here is what actually happens during inference — and why it explains latency, throughput, and per-token pricing.
An AWS blog post explains four deployment patterns for serving Unsloth-quantized models on AWS infrastructure, covering EC2, SageMaker AI endpoints, EKS, and ECS.
AWS details five new inference capabilities for SageMaker HyperPod, including multi-tier data capture, direct Hugging Face Hub deployment, NVMe model loading, Route 53 DNS, and pod-level IAM.
OpenAI introduces GPT-5.6, a frontier model it says delivers more intelligence per token, better performance per dollar, and additional capacity for demanding tasks.
Modern AI runs on a scarce resource: specialized compute. This feature unpacks why GPUs became the bottleneck, why everyone is building custom chips, and what it means for the balance of power in AI.
The race is no longer only about bigger models. Small language models that run cheaply — even on a phone — are one of AI's most important trends. Here is why, and what to watch.