The chip war just escalated. OpenAI’s Jalapeño beats Nvidia’s Rubin. Nvidia answered with Groq 3 LPX and a 15% price hike warning. Z.AI dropped GLM-5.3 Flash. Open weights, natively multimodal, near Claude Opus 4.8 at 1/10 the price. Reddit’s ChatGPT citations collapsed 86% in three days. A warning for anyone building on AI-search traffic.
👋 For the vibe-coders: Codex Resets this week
Shoutout to everyone running on Codex: there were 3 usage resets this week — August 24, 25, and 27 (the last one at 4:35 PM UTC). Resets get announced by @thsottiaux (OpenAI) on X, and the easiest way to track them is codex-resets.com — it’s counted 47 resets total, averaging ~7.5 days apart. Three in one week is well above average.
Today we have:
Featured Materials 🎟️
News of the week 🌍
Useful tools ⚒️
Weekly Guides 📕
AI Meme of the Week 🤡
AI Tweet of the Week 🐦
Bonus Materials 🎁
Your AI Product could be featured here!
Showcase your AI products, agents and models in front of 40k AI-native founders, creators and c-levels
Featured Materials 🎟️
Nvidia’s Double Move: Buys Hugging Face for $12.9B + Posts Record $96.2B Quarter 🎯
The chip war just escalated. OpenAI says its custom inference chip “Jalapeño”, built with Broadcom, taped out in 16 months on TSMC N3P and beats Nvidia’s Rubin on several efficiency benchmarks (Aug 25). Nvidia answered the same week: Groq 3 LPX (the inference accelerator from its $20B Groq acquisition) hit full production, with Nebius as the first cloud customer deploying it (Aug 24) — while Nvidia also warned server builders that AI system prices will rise more than 15% starting on shipments in early 2027 (Aug 22).
On August 26, Nvidia closed its $12.9 billion acquisition of Hugging Face — the largest acquisition in the company’s history and the biggest open-source AI platform deal ever. The next day, Nvidia reported Q2 FY2027 revenue of $96.2 billion, up 106% year-over-year, with data center revenue hitting $89 billion — also a record.
For scale: AM Intelligence (a Greenko subsidiary) placed a binding order for ~9,000 Vera Rubin systems — roughly $8B in planned AI infrastructure (Aug 26).
What Nvidia actually bought:
- Hugging Face hosts over 1 million models and serves as the default repository for open-weight AI development. Nvidia now controls the distribution layer for most open-source models — the same models that run on Nvidia GPUs.
- The deal includes Hugging Face’s enterprise platform and its 10,000+ paying customers, including Google, Meta, and Microsoft. Nvidia gains direct relationships with every major AI lab that builds on its hardware.
- Hugging Face’s open-source credibility was the key asset. Nvidia has been building Nemotron and NeMo, but owning Hugging Face means owning the community where developers discover, compare, and deploy models.
The quarter that funded it:
- Data center revenue: $89B, up 106% YoY, representing 92% of total revenue.
- Gross margin: 76.2%, up from 75.1% last quarter — pricing power remains intact.
- Nvidia’s market cap crossed $4.2 trillion on the news, adding $200B in a single day.
Nvidia didn’t just buy a platform. It bought the place where the next generation of AI developers learns what to build. The community that once defined open-source independence is now a division of the chip company that supplies almost every data center on earth.
Buying the library means owning the reader. Buying Hugging Face means Nvidia no longer needs to convince developers to use its tools — it just needs to keep the repository running and the GPUs humming.
Sources: TechCrunch | NVIDIA Newsroom |
GLM-5.3 Flash: Open Weights, One-Tenth the Price, Near Opus 4.8 ⚡
Z.AI released GLM-5.3 Flash on August 26 — the first natively multimodal model in the GLM series. At 320B total parameters with 18B active per forward pass, it scores within striking distance of Claude Opus 4.8 while costing roughly one-tenth the price.
What the numbers actually say:
MIT-licensed and open weights — the first frontier-class model from a Chinese lab released under permissive commercial terms. Deploy it, fine-tune it, fork it, sell it.
Native multimodal from the start — not a separate vision model bolted on. GLM-5.3 Flash handles images, video, and text in a single architecture. The demo shows real-time video understanding at 30 frames per second.
Benchmarks: On OpenRouter’s leaderboard, GLM-5.3 Flash trails Claude Opus 4.8 by 2.3 points on the Intelligence Index — 87.4 vs 89.7 — at $0.42 per million input tokens vs Opus 4.8 at $3.75. That is not a gap. That is a rounding error with a decimal shift.
Z.ai is expected to publish the weights following a two-week safety hold; licensing terms haven’t been announced yet.
Why this matters right now:
DeepSeek V4 Pro shipped at $0.87 per million output two weeks ago. GLM-5.3 Flash ships at $0.42 input and $0.84 output — the same price class, but natively multimodal and open weights. The open-weight tier just got more capable, more multimodal, and cheaper — all at once.
The gap between “frontier” and “good enough” is now measured in cents, not capabilities. Chinese labs just proved that open weights can ship multimodal at 1/10 the price of the best proprietary model — and MIT-license it for commercial use.
Source: Z.AI
Reddit’s ChatGPT Citation Share Collapsed — 3.83% to 0.52% in Three Days 📉
Reddit’s share of ChatGPT citations fell from 3.83% to 0.52% in just three days after an unannounced change to ChatGPT’s retrieval system (Aug 26). For creators and publishers who had started treating AI-search citations as a real traffic channel, this is the first big proof of how fragile that distribution actually is.
What actually happened:
3.83% → 0.52% — a drop of nearly 90% in citation share. The change was unannounced, undocumented, and rolled out silently.
Reddit was the canary in the coal mine. If Reddit — one of the most-cited domains in AI training data — can lose almost all its citation share overnight, any publisher building a strategy around AI-search traffic is building on sand.
The trigger: an unannounced retrieval change to ChatGPT’s citation system, likely shifting toward newer or more authoritative sources. OpenAI hasn’t commented on the change.
Why this matters for creators and publishers:
AI-search citations had become a real traffic channel. Publishers optimized for them. Some built entire strategies around “being cited by ChatGPT.” This collapse shows that distribution is not earned — it’s leased. And the lease can be revoked without notice.
If a 90% drop can happen to Reddit in three days, it can happen to anyone. AI-search traffic is not a channel — it’s a dependency. And dependencies that change without warning are not distribution; they are rent.
Source: Promptwatch
Keep your mailbox updated with practical knowledge & key news from the AI industry!
News of the week 🌍
Emerald AI Raises $150M Series A at $1.05B Valuation for Grid-Flexible Data Centers⚡ — Emerald AI announced a $150 million Series A on August 25, led by Nvidia, Samsung, and Siemens, valuing the company at $1.05 billion. The startup builds software that shifts data center workloads in real time based on power grid conditions — reducing energy costs by up to 40% and avoiding peak demand charges. The funding signals that the AI compute bottleneck is no longer just chips: it is the electricity to run them, and investors are betting on software to solve the power problem before the grid does.
Nvidia Signs $6B Licensing Deal with Poolside to Build Open-Weight Models💰 — Nvidia signed a $6 billion AI model licensing agreement with startup Poolside on August 23, the Wall Street Journal reported. Nvidia will also invest an additional $1 billion, bringing Poolside’s pre-money valuation to $12 billion, and will offer jobs to about 100 Poolside employees to contribute to the Nemotron project. The goal: build world-class open-weight AI models to compete with China’s DeepSeek and Kimi.
SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Scale⚡ — On August 24, NVIDIA announced that SpaceXAI will deploy its new NVIDIA Vera CPU — the first CPU purpose-built for AI agents — to accelerate agentic AI workloads behind Grok. Vera delivers up to 1.8x faster task completion than x86 CPUs. SpaceXAI plans to expand its infrastructure on NVIDIA Vera Rubin toward megawatt-scale compute capacity, and even extend it to orbit with first-generation Starmind AI satellites.
GLM-5.3 Flash Stealth-Launched on OpenRouter as ‘Ox Alpha’ and Topped Token Charts ⚡ — An anonymous model called “Ox Alpha” appeared on OpenRouter on August 20 with no announcement, no documentation, and a free 1M context window. By August 23, it had climbed to the top of OpenRouter’s weekly token consumption chart, outpacing every major frontier model in usage volume. Z.ai later confirmed that Ox Alpha was actually GLM-5.3 Flash — the natively multimodal model they officially announced on August 26. The stealth launch was an unannounced stress test that accidentally became the week’s biggest model story.
SuperGrok Heavy Quietly Dropped ‘Near-Unlimited Usage’ From Its Pricing Page 📝 — xAI removed the “near-unlimited usage” promise from SuperGrok Heavy’s $300/month pricing page, replacing it with “highest usage limits.” The change was spotted and discussed by users on August 25. The distinction matters: “near-unlimited” was a promise about the experience; “highest limits” is just a comparison to cheaper tiers. Users who signed up under the original promise are now living with a materially different product description — and materially different limits (3–5× the capacity of SuperGrok for 10× the price).
Salesforce: Companies Now Run 13 AI Agents on Average — Up From 5 in Early 2025 📊 — Salesforce’s second Agentic Enterprise Index, published August 26, found that the average number of AI agents deployed per organization nearly tripled from 5 in early 2025 to 13 by April 2026. Agent creation time fell 53% to an average of 1.9 days. In customer service, 7 in 10 sessions are now handled autonomously — and critically, escalations to humans stayed steady, confirming that the automation is real deflection rather than just hiding the contact button. The data is based on production telemetry from 400 businesses, not a survey.
Useful tools ⚒️
⭐ Skild S1 — a new robotics foundation model that lets robots learn and execute physical tasks up to 10 minutes long from a single human video demonstration, with no fine-tuning required. Built on a massive manipulation dataset, S1 adapts on the fly with no retraining, achieving an average success rate of 66% per step for unseen tasks — significantly outperforming traditional VLA models at 9%. Released Aug 26.
Pinecone Nexus — a knowledge engine for agents that reached general availability Aug 6. Nexus compiles your enterprise data into a queryable layer once, then serves grounded, cited answers to agents on every call. In benchmarks, GPT‑5.5 with Nexus held accuracy at 77% less cost per task; GPT‑5.2 gained 12% more accuracy and saw 80% cost reduction. Both cut tool calls and model calls roughly in half.
Cloudflare Browser + Wallet for agents — a genuinely new product: Cloudflare launched Cloudflare Wallets and cloudflare.pay on Aug 4 to give AI agents a verifiable identity and controlled spending limits. Alongside, Kitesurf — a cloud‑hosted, agent‑first browser that lets agents navigate websites and complete browser‑based tasks without teams managing their own browser infrastructure. Together, they give autonomous agents an identity, a wallet, and a browser to transact online.
IdeaHunter — a fresh launch for solo founders that surfaces demand‑backed product ideas from real market signals: complaints, feature requests, job posts, app reviews, and open‑source activity across Reddit, Hacker News, Product Hunt, GitHub, Upwork, and more. Each opportunity comes with a specific buyer, MVP scope, monetization path, and build‑time estimate — grounded evidence, not generic AI brainstorming. Free plan with 3 weekly deep dives; Pro at $9.99/month. Launched Aug 23.
Diet Claude — Never hit Claude’s rate limits again. Diet Claude is a free Chrome extension that shows your remaining Claude usage at a glance — session limits, weekly quotas, reset timers — and automatically compresses long conversations to stay within token budgets. The compression feature uses a custom algorithm that reduces context without losing meaning, letting you squeeze more work out of each session.
Share this post with friends, especially those interested in AI!
Weekly Guides 📕
The New Creator Economy: From AI Clones to 100M-View Reels — Daniil’s deep dive on six creators building businesses with AI: Kwebbelkop’s failed clone experiment, Coffee&Pages’ 100M-view Instagram Reels, Kavan the Kid’s physical-mask-to-AI workflow, Granny Spills’ synthetic influencer, and YouTube creators launching games without code. Essential reading if you’re building a creator business in 2026 — the framework is more durable than any tool stack.
GLM-5.3 Flash — Getting Started Guide — Official documentation from Z.AI covering model architecture, API setup, multimodal capabilities, and performance benchmarks. Includes sample code for text, image, and video inputs. Essential reading before migrating any workflow to the new model — especially if you’re comparing it to DeepSeek V4 Pro or Claude Opus 4.8.
How to Build AI Agents for Total Beginners (2026) — A clean, no‑fluff entry point from The Neuron that walks you through agent fundamentals, tool calling, and memory without assuming any prior AI experience. Perfect if you’re starting from zero and want to understand what agents actually do before you build one.
AI Builder Week 2026 — 10+ free online workshops from Product School running this week, covering hands-on sessions on agent workflows, prompt engineering, and shipping AI products. A solid alternative if you prefer building over reading — and it won’t cost you anything to join.
AI Meme of the Week 🤡
AI Tweet of the Week 🐦
Bonus Materials 🎁
Inherent (DeepMind alumni) — AI Teammate Outperforms Anthropic & OpenAI at Replicating Research — Founded by DeepMind alumni, Inherent claims its AI teammate can independently replicate research workflows — scoping experiments, writing code, running analyses, and producing draft papers — outperforming Anthropic and OpenAI agents on internal benchmarks. The company emerged from stealth on August 22, announcing a $200M round led by Sequoia. The claim is unverified, but the team’s lineage makes it worth watching.
Techno Machine in One HTML File — a developer built a full techno drum machine into a single HTML file with “verifiable renders.” 314 points and 151 comments on Hacker News this week — pure joy for anyone who loves minimalist engineering flexes.
CarWatch — turn your car into a chat-room agent: Raspberry Pi 5 + dashcam + a local Qwen model, fully offline. A great “AI for the garage” project to try over a weekend.
Anthropic — Model Hardware Standard (MHS) Research Preview — Anthropic’s research preview on a new standard for describing model capabilities in hardware terms. The proposal aims to standardize how AI models report their compute requirements, performance characteristics, and hardware compatibility — making it easier to match models to the right infrastructure. If adopted, it could reduce the chaos of deploying models across different GPU types and cloud providers.
If you missed our previous updates, don’t worry, here they are:
Agents Go Rogue, SpaceX Burns $18B, Open Weights Win | Weekly Digest
Your take: Nvidia owns the world’s largest open-source AI platform and just posted a record quarter where data center revenue hit $89B — 92% of total revenue. Z.AI shipped a MIT-licensed multimodal model that nearly matches Claude Opus 4.8 at 1/10 the price. Australia banned AI music from its charts after an AI-generated Madonna cover hit #1 with 48M streams. Is the AI industry consolidating around Nvidia faster than anyone can compete — or are open-weight models from China the real story here? Drop it in the comments 👇








