25 articles tagged with “multimodal”.
AI NewsAWS explains how to build a WhatsApp ordering assistant on Amazon Bedrock AgentCore with Amazon Nova 2, accepting text, voice notes, and real-time calls on one number with shared memory across channels.
AI NewsHugging Face has unveiled NeoMME, a new encoder designed for efficient multimodal and multilingual processing.
AI NewsAWS describes building a generative AI support operations platform that converts training videos into structured SOPs, applies RAG to guide ticket resolution, and predicts SLA risk to prioritize work.
AI NewsGoogle's Gemini Notebook can now import purchased Google Play Books via a new Expert Intelligence feature, letting users query titles and generate plans, infographics, and AI podcasts from their contents.
AI NewsGoogle's new Gemini 3.5 Transcribe model adds transcription that detects specialized jargon, supports more than 85 languages, and edits out filler words like 'ums' and 'ahs'.
AI NewsAWS outlines a multi-agent document classification approach on Amazon Bedrock that combines text analysis with Claude Haiku 4.5 and visual similarity search using Titan Multimodal Embeddings to classify insurance documents.
AI NewsTwitch has added an option letting users exclude their content from future training of Amazon's generative AI models, while other AI-supported features continue to work.
AI NewsGoogle's Gemini app has reached 1 billion users, with the company reporting 63% of users engage via voice and more than 150 million images generated daily.
AI NewsGoogle DeepMind says Gemini Robotics 2 can control a humanoid robot's entire body, expanding beyond the previous model's upper-body focus to full-body motion.
AI NewsGoogle DeepMind introduced Gemini Robotics ER 2, a model aimed at helping robots reason, collaborate, and solve real-world tasks through improved video understanding and tool orchestration.
AI NewsAnthropic is bringing Claude's voice mode to its Opus and Sonnet models, moving beyond the faster Haiku model, and extending voice into apps like Gmail, Slack, and Canva.
AI NewsSamsung has shown its upcoming AI smart glasses in person, unveiling two new designs and initial specs including 9-hour battery life, with a fall launch planned.
AI NewsMeta unveiled Content Seal, an invisible watermarking tool that flags images from its AI model, but the tool is seen as less accessible and reliable than existing options like SynthID and C2PA.
AI NewsAdobe is adding a generative AI suite called AI Playground to its experimental Project Indigo camera app, offered as an opt-out experiment to a small percentage of users.
AI NewsNVIDIA used SIGGRAPH to showcase graphics and simulation advances driven by agentic and physical AI, spanning media, content creation and robotics.
AI NewsAmazon Bedrock now offers access to Grok 4.3, positioned for agentic and enterprise workloads with features such as configurable reasoning effort, tool calling, and multi-turn conversations.
AI NewsGoogle Vids is introducing personalized AI avatars that let users appear in their own AI videos, along with Gemini Omni-powered generation and editing from prompts and reference images.
AI NewsGoogle is rebranding NotebookLM as Gemini Notebook. The app remains standalone but will integrate more closely with Gemini and Google Search.
AI NewsAWS describes a Computer Vision MCP Server built with Amazon Bedrock, offering a standardized interface for AI systems to process visual information and make decisions.
GuidesSpeech recognition went from a brittle research problem to a solved-ish commodity in about three years. Here is the pipeline that made it work — and where it still breaks.
AI NewsWaze is gaining four new updates, two of which use Google's Gemini assistant, including conversational voice reporting and a new Destination Search feature.
GuidesSearching photos by typing a description, or matching an image to a caption, relies on one clever idea: putting text and images into the same mathematical space. Here is how CLIP made that work.
AI NewsThe leading AI models no longer just read and write text — they see images and hear audio too. Here is what 'multimodal' means and why it has become the default expectation.
AI ToolsStable Diffusion made high-quality text-to-image generation open and runnable on consumer GPUs, sparking a vast ecosystem of tools and fine-tunes. Here is what it is and why it mattered.
ResearchCLIP learned to link pictures and language by studying hundreds of millions of image–caption pairs from the web. It quietly became the foundation for image search and generation.