AI Model News

Catch up on the latest news on AI models

Why GLM-5.3-Flash and Qwen3.8-Flash-Next Landed on the Same Design

Two Chinese AI labs, Z.ai and Alibaba's Qwen team, built new open AI models this week without talking to each other. Their models, GLM-5.3-Flash and Qwen3.8-Flash-Next, ended up using nearly the same design tricks to run faster and cheaper, though they disagreed on how to handle position information in text.

MarkTechPost — August 28, 2026

Google's Gemini 3.5 Transcribe Hits 2.6% Word Error Rate in Tests

Google launched Gemini 3.5 Transcribe, a new speech-to-text model that reports a 2.6% error rate on recorded audio and 4.0% on live streaming across more than 85 languages. It comes as two separate tools, one for recordings with speaker labels and word timestamps, and one for real-time streaming that is faster but skips those extra features.

MarkTechPost — August 28, 2026

Gemini Omni 1.1 Flash adds more control for video generation

Gemini Omni 1.1 Flash gives developers more power to build videos, with longer scene extensions, set start and end frames, and fast 360p drafts. It also adds 4K upscaling and video references, and is now live in Google AI Studio, Agent Platform, and Google Flow.

DeepMind — August 27, 2026

Z.ai's GLM-5.3-Flash Brings 1M-Token Context to a 320B MoE Model

Z.ai has launched GLM-5.3-Flash, a huge but efficient multimodal AI model that can read text, images, and video, and it handles up to 1 million tokens of context at once. It costs far less than older models, nearly matches Claude Opus 4.8 on coding tasks, and can be run by companies on their own powerful servers or used through Z.ai's online service.

MarkTechPost — August 26, 2026

Google launches Gemini 3.5 Transcribe for smarter speech-to-text

Google's new Gemini 3.5 Transcribe model turns speech into clean, accurate text, even fixing self-corrections and cutting out filler words like "um" and "ah". It works fast in over 85 languages and is already built into tools like Gboard, the Gemini app, and Chrome, with more error correction and speed than the older Chirp 3 model.

DeepMind — August 26, 2026

Qwen3.8-Flash-Next Debuts as Early Preview of Qwen4's Architecture

Alibaba's Qwen team released Qwen3.8-Flash-Next, a huge open-weight model that only turns on 6 billion of its 180 billion parameters for each word it processes, making it cheap to run while still handling text, images, and code. It trains for about one-ninth the cost of the earlier Qwen3.7-Plus model, scores well on coding and agent tests, and points toward the design Alibaba plans to use for its next big model, Qwen4.

MarkTechPost — August 26, 2026

OpenAI's new Jalapeño chip poses growing threat to Nvidia's margins

OpenAI has revealed its first AI chip, called Jalapeño, and analysts say it could hurt Nvidia's profits. The chip matches or beats Nvidia's top chips at running AI systems, joining a wave of custom chips from Google, Amazon, and Meta that threaten Nvidia's grip on the market.

CNBC — August 26, 2026