Rapidly scaling online storage to serve over 1 billion ChatGPT users
OpenAI details how its Habitat storage system grew from a Python library into a globally distributed platform supporting 1 billion ChatGPT users and 22 million requests per second.
96 articles tagged with “cloud”.
OpenAI details how its Habitat storage system grew from a Python library into a globally distributed platform supporting 1 billion ChatGPT users and 22 million requests per second.
Amazon SageMaker Inference introduces prefix-aware routing, directing requests with shared prompt prefixes to the same instance to keep KV caches warm and reduce latency.
SageMaker HyperPod now supports model caching for inference, pre-loading model weights and container images onto cluster nodes so pods read from local NVMe instead of downloading over the network.
NVIDIA's GeForce NOW cloud gaming service adds nine new titles this week, including WARDOGS at early-access launch, a Valheim 1.0 update, and Bus Simulator 27.
A July 2026 transmission line fault in Ashburn, Virginia dropped more than 3 gigawatts of data center load in seconds, echoing an earlier incident and pointing to grid architecture challenges for AI infrastructure.
AWS details how to deploy the open-weight Qwen3.8-2.4T-A95B model on Amazon SageMaker HyperPod with vLLM, covering cluster setup, quantization, and an OpenAI-compatible endpoint.
AWS summarizes its August 2026 releases for AI builders spanning Amazon Bedrock, Bedrock AgentCore, and Strands, including million-token context, cross-Region inference, and long-running agents.
OpenAI's GPT-6 Astra is generally available on Amazon Bedrock, offering deeper reasoning on Bedrock's inference engine built for performance, security, and scale.
Pathway's Baby Dragon Hatchling is a brain-inspired, post-transformer architecture that reasons in latent space and is developed and scaled on Amazon SageMaker HyperPod.
Amazon SageMaker Feature Store introduced the UpdateRecord API, enabling updates to one or more feature values in a single call without reading or rewriting the full record.
AWS extends MLflow and SageMaker AI Model Registry sync to two cross-account governance patterns: a hub-and-spoke model centralized with AWS RAM and a hybrid model that isolates development accounts.
AWS outlines how to build memory lifecycle policies for Amazon Bedrock AgentCore, scoring, consolidating, and pruning agent memories through a nightly Step Functions workflow with a deployable CDK stack.
AWS introduced HyperPod InstantStart, an open source control plane combining Amazon EKS with SageMaker HyperPod to run guarded operations via a web interface and an AI agent.
A guide to customizing Amazon Bedrock knowledge bases for large, complex documents by pairing Amazon Textract text extraction with generative AI to query documents like utility bills at scale.
AWS details how to run a customer-operated LiteLLM gateway on Amazon ECS with Fargate, link it to an OpenAI model on Bedrock, and route Codex requests with scoped identities, budgets, and telemetry.
Amazon Bedrock now offers Australian teams access to OpenAI GPT-5.6 Sol, Terra, and Luna models with global cross-Region inference from the Sydney and Melbourne Regions.
University Startups and AWS partner g/d/n/a built Trinity, a serverless multi-agent AI on Amazon Bedrock that creates IDEA-aligned transition plans for students with disabilities.
Google has announced several AI updates from August 2026, highlighting advancements in enterprise solutions and automation. The updates focus on enhancing cloud capabilities and AI integration.
AWS has made Claude Fable 5.1 available on Amazon Bedrock and the Claude Platform, alongside Enterprise Frontier Safeguards for keeping data in a controlled cloud environment.
Nvidia is launching its DLSS 5 AI upscaling technology this week for RTX 50-series GPUs and GeForce Now, with NBA 2K27 as the only supported game at release.
An AWS tutorial covers deploying and hosting an MCP server in AgentCore Runtime and integrating it with Amazon Quick to enable reusable AI tools and agents.
New data center systems reportedly gain efficiency through smarter traffic control rather than relying solely on more processor cycles, extending Nvidia's advantage beyond the GPU.
Lambda secured $1B in private debt to buy Nvidia AI chips and lease them to Microsoft, adding to a series of loans that reflect the high cost of the AI boom.
Sporting goods retailer Decathlon uses Chronos-2 on AWS to forecast weekly demand across tens of thousands of products, reporting accuracy gains of 11-15 points and low-cost CPU inference.
Amazon plans to add another 2 million Nvidia GPU chips to its data centers over the next two years, citing surging demand and an expanded partnership with Nvidia.
NVIDIA expands its NVLink Fusion platform with NVHBM custom high-bandwidth memory, targeting the infrastructure demands of AI agents and trillion-parameter workloads.
Amazon OpenSearch Service now supports MCP Apps, enabling AI agents to move from alert to root cause in one conversation while returning interactive visualizations inline.
AWS outlines a governed weekly reporting workflow combining Amazon Quick Desktop and Amazon FSx for NetApp ONTAP, drafting cited reports and Slack summaries with human review before sharing.
SageMaker HyperPod now offers managed Ray support on Amazon EKS, letting users create and monitor Ray clusters, connect notebooks, and run distributed training and accelerated inference.
AWS describes how to build a customizable knowledge management system that delivers institutional knowledge via a voice-first AI avatar, using Amazon Bedrock Knowledge Bases for RAG.
AWS unveiled Agentic Resource Discovery (ARD), an open specification, and AWS Agent Registry, a centralized searchable catalog for agents, tools, and skills across environments.
NVIDIA describes AI factories as full-scale infrastructure whose value is measured by delivered output metrics, and considers how custom XPUs fit into that model.
Nvidia has partnered with data center developer Cloverleaf, extending its ongoing investment in data center development.
Amazon Bedrock now supports OpenAI GPT-5.6 models with cross-Region inference in more than 25 AWS Regions, routing requests for higher throughput via US geographic and global inference profiles.
AWS describes a multi-agent framework built on Amazon Bedrock AgentCore that automates enterprise cloud migrations end to end, cutting IaC development from weeks to minutes.
An AWS post outlines vector search capabilities integrated into its existing databases and storage services, covering six purpose-built services and guidance for selecting an engine.
An AWS blog post describes three serverless patterns for asynchronously calling Amazon Bedrock AgentCore agents from AWS Step Functions pipelines to avoid idle compute costs.
TerraPower's nuclear power plant reportedly holds a strategic advantage over competitors as it pursues deals to power AI data centers.
This article discusses the memory demands of AI agents, emphasizing the importance of optimizing memory usage for enhanced performance.
Jumio built a centralized, real-time feature store on AWS delivering sub-100ms feature serving for fraud detection while saving roughly $120,000 annually.
Cybersecurity SaaS provider Axonius used Amazon Bedrock AgentCore to deploy fully isolated, multi-tenant AI agents across hundreds of customer environments.
NVIDIA Nemotron 3.5 Lightning, an open 30B Mixture-of-Experts model with 3B active parameters, is now available in Amazon SageMaker JumpStart for high-volume agentic workloads.
NVIDIA frames AI factories as the defining infrastructure of the AI era, requiring advanced chips, packaging, memory, networking, plus land and power.
AWS describes building multi-agent workflows using OpenAI-compatible SageMaker AI endpoints and Bedrock AgentCore runtime, with each agent using a model suited to its task.
A new forecast suggests natural gas prices could triple in some U.S. regions, potentially leaving hyperscalers with large bills to power their AI data centers.
Sheets canvas enables users to create interactive dashboards, study trackers, and seating charts from spreadsheet data using easy prompts.
AWS details how Amazon Bedrock AgentCore Observability can track AI agents running outside AWS, routing traces, span metrics, and token usage to one dashboard.
NVIDIA's GeForce NOW releases its native Linux app out of beta, adds cloud optimizations for Frame Generation, and raises frame rates for Performance members.
This AWS guide explains how to set up CUR 2.0 with IAM principal data and use Amazon Athena and CUDOS dashboards to track and analyze Amazon Bedrock costs across an organization.
Google introduced the Pixel 11 lineup and a new competitor to AirTag, alongside exciting Gemini features at the Made by Google 2026 event.
An AWS blog post details building a tiered KV cache on Amazon SageMaker HyperPod with Curvine, extending the cache into a shared NVMe pool so replicas reuse cache on cost-efficient instances.
OpenAI's specialized cyber defense models, Daybreak Red and Daybreak Blue, are now available on Amazon Bedrock for eligible customers, running with zero-operator access enforced at the chip.
ONESTRUCTION, advised by the AWS Generative AI Innovation Center, built Ishigaki-IDS, a foundation model for construction and BIM workflows using synthetic data and a three-stage training pipeline on Amazon EC2.
AWS presents a reference deployment for Claude apps gateway, a self-hosted governance layer connecting Claude Code and Claude Desktop to Amazon Bedrock or Claude Platform on AWS.
NVIDIA outlines why scaling AI compute requires rethinking power delivery, noting that the challenge lies in how electricity moves from the grid to the GPU, not just total wattage.
AWS describes how the SageMaker AI Spaces add-on for Amazon EKS runs managed JupyterLab and Code Editor environments on an existing cluster, with browser and VS Code access.
nOps moved its Clara FinOps AI agent from a self-managed EKS stack to Amazon Bedrock AgentCore, reducing time-to-production by 75% while improving response quality.
Firebird, an emerging AI cloud, has launched what it describes as the CIS region's largest AI factory in Armenia, using NVIDIA accelerated computing and Dell high-performance infrastructure.
TReNDS, a Georgia State University research center, built an agentic AI pipeline on Amazon Bedrock and the Strands Agents SDK that investigates production errors in real time.
Cloudflare has introduced Kitesurf, a cloud-hosted browser designed for AI agents rather than human users, aimed at helping developers build browser-based automation more efficiently.
Mobileye deployed an AI agentic support solution on Amazon Bedrock AgentCore, using a hybrid architecture to connect on-premises systems with AWS cloud while maintaining governance and security.
AWS details a secure MCP bridge that connects a cloud-hosted Bedrock AgentCore agent to local MCP servers by tunneling signed messages over an existing WebSocket connection.
Amazon Bedrock AgentCore harness is generally available and can be added as an agent step in n8n workflows through a new open-source community node.
Amazon Bedrock now offers Web Search as a generally available built-in tool, letting foundation models ground responses in current web knowledge without third-party vendors or external API orchestration.
NVIDIA highlights how growing AI demands push datasets and context windows past system memory, requiring efficient, secure storage architectures rather than added capacity alone.
AWS is enabling the vibe-coding tool Superblocks to be embedded into the private clouds of AWS customers, described as a step toward decoupling apps from models.
Formula 1 worked with AWS to build a Data Accelerator using agentic AI on Amazon Bedrock AgentCore, reducing data source onboarding from up to eight weeks to roughly 40 minutes.
OpenAI GPT-5.6 Sol, Terra, and Luna are generally available on Amazon Bedrock, now with explicit prompt caching that lets users choose which prompt parts to cache and reuse.
Nscale, a British AI neocloud, is acquiring Anyscale, a software startup that helps companies scale AI workloads across data centers and servers.
Meta CEO Mark Zuckerberg told investors on the company's second-quarter earnings call that its enterprise AI opportunity spans agents, APIs, compute, and internal software.
An AWS guide covering Private Key JWT client authentication in Amazon Bedrock AgentCore Identity, including supported grant flows and setup steps for signing keys and credential providers.
The Model Context Protocol released its 2026-07-28 spec, its largest revision yet, and AWS details how to enable it on AgentCore Gateway via a single UpdateGateway call.
Google AI has unveiled new features for Managed Agents in the Gemini API, aimed at enabling developers to create reliable agents for production use.
Grid operators may impose temporary power cuts on data centers to prevent blackouts on the largest US grid, as rapid data center construction strains power supply.
Recursive Superintelligence has agreed a $410M compute deal with Amazon, directing spending toward compute rather than headcount as it works to automate its own product development.
An AWS blog post presents task-aware knowledge compression (TAKC), which pre-compresses knowledge bases into task-specific representations across fidelity tiers to handle analytical tasks spanning many documents.
A close call from a fallen power line in Northern Virginia highlighted weaknesses in how data centers respond to grid disruptions, with a look at potential fixes.
Three OpenAI GPT-5.6 models—Sol, Terra, and Luna—are now generally available on Amazon Bedrock, with support for the Responses API, prompt caching, and the Codex coding agent.
An AWS tutorial explains building multi-Region carrier performance dashboards in Amazon QuickSight with Highcharts custom visualizations while maintaining data sovereignty across Regions.
A projection indicates data centers could quadruple their electricity consumption by 2035, with new facilities built through 2033 potentially using as much power as India does today.
NVIDIA's Vera Rubin is moving into production, with NVL72 racks deployed at major cloud partners and backed by a supply chain spanning over 350 factory sites in 30 countries.
Couchbase adopted Amazon Bedrock with Anthropic's Claude models to power Capella iQ, using a multi-model architecture and reporting operational benefits in production.
Hugging Face has launched Cosmos 3 Edge, a new tool aimed at developers to enhance AI model development and deployment on edge devices.
An AWS blog post outlines how Smartsheet built a remote MCP server on AWS, detailing the underlying infrastructure including security, governance, scaling, deployment, and AI-specific optimizations.
AWS details how to build a telephony AI host that answers calls and takes restaurant orders using Amazon Bedrock AgentCore, Nova 2 Sonic, and the Model Context Protocol.
NVIDIA's GeForce NOW cloud gaming service is adding Onimusha: Way of the Sword at launch alongside new titles, and is now publicly available in India after moving out of beta.
Built Technologies worked with AWS teams to create an AI-powered document processing engine for real estate finance, cutting multi-day workflows to minutes across hundreds of document types.
AWS presents a solution for centralized monitoring of SageMaker Pipelines across AWS accounts and Regions using CloudWatch custom dashboards, backed by a CDK example on GitHub.
An AWS blog post outlines a cloud-based UX testing platform that uses Amazon Nova Act to auto-generate test scenarios from documentation and run user flows in parallel at scale.
OpenAI's GPT-5.6 Sol, Terra, and Luna models are now generally available on Amazon Bedrock, accessible through its inference engine.
AWS details how to implement on-behalf-of (OBO) token exchange for multi-tenant agents using Bedrock AgentCore Gateway, including JWT claim transformations and audience binding.
AWS has launched a UI in Amazon SageMaker AI Studio that guides users to optimized generative AI inference configurations without deep infrastructure expertise.
The Verge's Stepback newsletter looks at rising community resistance to AI data centers, tracing current disputes back to early opposition to an Apple data center in Ireland.
An AWS blog post explains four deployment patterns for serving Unsloth-quantized models on AWS infrastructure, covering EC2, SageMaker AI endpoints, EKS, and ECS.
Microsoft's 2026 sustainability report says its carbon emissions rose 25 percent in 2025, reaching 34 million metric tons, driven primarily by datacenter expansion.
AWS details five new inference capabilities for SageMaker HyperPod, including multi-tier data capture, direct Hugging Face Hub deployment, NVMe model loading, Route 53 DNS, and pod-level IAM.