Latest / Tech Talks With Kinsoft / Last Week in Tech – Anthropic vs Alibaba, OpenAI's Own Chip, and an Emergency Cisco Patch
Transcript
- 0:00Imagine finding out that the biggest threat to
- 0:03a multi -billion dollar AI isn't some highly
- 0:06sophisticated nation state super virus. Right.
- 0:09But instead it's just a quiet army of like 25
- 0:12,000 bots asking it to write cunt for pennies.
- 0:14Yeah, it's wild. Or that the key to breaking
- 0:19a trillion dollar hardware monopoly is essentially
- 0:21just a software universal translator. Welcome
- 0:25to this custom deep dive. Today our mission for
- 0:28you listening is to really look underneath the
- 0:30hood of the AI applications you use every single
- 0:33day and well... examine this massive structural
- 0:36shift that's happening right now. Yeah. We're
- 0:38working from a really fascinating, incredibly
- 0:40dense dispatch today from a technology update.
- 0:43Right. Tech Talks with Kinsoft. Exactly. From
- 0:45June 29th, 2026. Yeah. And I mean, this source
- 0:49covers just one week in the tech industry. Yeah.
- 0:52But it acts like this perfect time capsule. So,
- 0:54okay, let's unpack this. Because reading through
- 0:56this dispatch, you realize really quickly that
- 0:57the, you know, the experimental novelty phase
- 1:01of AI, it's completely over. Oh, completely.
- 1:03It's done. We are looking at high stakes battlegrounds
- 1:06over intellectual property, massive hardware
- 1:10monopolies, and basically the dawn of automated
- 1:14cyber warfare. It is a really dense week of news.
- 1:18And I think the through line connecting all these
- 1:20seemingly separate events is really just one
- 1:23word, which is infrastructure. Infrastructure.
- 1:25Right. Yeah. Everything we're going to examine
- 1:27today really demonstrates how the very foundation
- 1:31of artificial intelligence is just being violently
- 1:33shaken up and rebuilt. Yeah, rebuilt from the
- 1:36ground up. Well, we are seeing a complete reconstruction
- 1:38happening on every single level. You know, how
- 1:41the underlying training data is safeguarded,
- 1:43how the physical silicon itself is manufactured,
- 1:46and how the entire security architecture of our
- 1:49networks is fundamentally changing. Well, I want
- 1:52to start with the intelligence itself. Because
- 1:54the Kinshoff dispatch kicks off with this wild
- 1:55story about a new kind of IP theft called distillation.
- 2:00Yes. Distillation. But before we get into the
- 2:02mechanics of this, and I want to make something
- 2:04explicitly clear to you listening, we are impartially
- 2:08reporting the contents of this source document
- 2:10today. Right. That's a very important distinction
- 2:12to make. Yeah, because the dispatch details accusations
- 2:16made by one company against another. You know,
- 2:20we are not endorsing these viewpoints and we
- 2:22aren't taking sides here. We are simply conveying
- 2:25the claims and the underlying technology that's
- 2:27described in the source material. Absolutely.
- 2:29And it's an important distinction because, well,
- 2:31the claims are explosive. On June 24th, 2026,
- 2:34the dispatch notes that Anthropic actually went
- 2:37to a U .S. Senate committee. Which is a big deal
- 2:40on its own. Right. And they presented a document
- 2:43accusing operators tied to Alibaba's Quen Lab
- 2:46of executing the largest known distillation attack
- 2:49on their clawed models. The numbers they presented
- 2:52to lawmakers are just, they're staggering. I
- 2:55mean, they alleged roughly 25 ,000 fraudulent
- 2:57accounts were created. 25 ,000. Just to bypass
- 3:00API rate limits. And through those accounts,
- 3:03these operators supposedly ran nearly 29 million
- 3:07conversations. Over the course of about six weeks,
- 3:11I think. Yeah, six weeks. And the goal here was
- 3:13to cheaply copy, or as they call it, distill
- 3:17Claude's capabilities, specifically targeting
- 3:20its coding, reasoning, and cybersecurity skills
- 3:23to train a rival model. Right. Now, Alibaba completely
- 3:27denies this, obviously, and as the source notes,
- 3:30it hasn't been independently verified. Putting
- 3:33the geopolitical drama aside for a second, let's
- 3:36talk about the mechanics here. Let's do it. Because
- 3:38I understand using bots to scrape data, right?
- 3:40That makes sense. But wait, if you're just copying
- 3:43the outputs of a model, how does a smaller competing
- 3:47model actually learn reasoning? That's a great
- 3:49question. Isn't it just like parroting answers
- 3:51without understanding the underlying logic? Well,
- 3:54what's fascinating here is the underlying mechanism
- 3:56of how synthetic data transfer works in modern
- 3:59neural networks. You aren't just scraping the
- 4:01final answers. When an operator runs a distillation
- 4:03attack on this massive scale, they're actually
- 4:06prompting the target model to generate what we
- 4:09call chain of thought output. Chain of thought,
- 4:11right. Right. They are asking the model to show
- 4:13its work step by step before arriving at the
- 4:15solution. Ah, I see. So instead of just saying,
- 4:18you know, write this final line of code. They're
- 4:21asking it to explain why it chose a specific
- 4:24variable or a specific algorithm in the first
- 4:27place. Precisely. To use an analogy, think of
- 4:30the frontier model, in this case, Claude, as
- 4:33a master teacher. Okay. A master teacher. Yeah.
- 4:35A teacher who has already spent years reading
- 4:37every single textbook in the world. A distillation
- 4:40attack is essentially like forcing that master
- 4:43teacher to sit down and write 29 million highly
- 4:47detailed step -by -step study guides. Wow. Okay,
- 4:51so then the smaller student model. The student
- 4:53model doesn't need to read the entire library
- 4:55of human knowledge anymore. Right. It just reads
- 4:58the study guides. Exactly. It trains directly
- 5:00on those pristine, highly structured study guides.
- 5:03It learns the logical pathways, the reasoning
- 5:06itself, without the massive computational overhead
- 5:08of having to figure it all out from scratch.
- 5:10That is just wild. I mean, it completely bypasses
- 5:13the billions of dollars and, what, months of
- 5:16continuous data center compute required for original
- 5:18training? Oh, absolutely. You are literally weaponizing
- 5:22your competitor's multi -billion dollar machine
- 5:25to train your own machine. Which is exactly why
- 5:28this is considered asymmetrical warfare. Yeah.
- 5:31The cost to defend against this is incredibly
- 5:33high. Because think about it, to the frontier
- 5:35model servers, this attack just looks like normal,
- 5:39albeit heavy, user traffic. Right, because it's
- 5:42distributed across thousands of Sybil accounts.
- 5:45Just looks like a bunch of users chatting. Exactly.
- 5:48And taking this to a U .S. Senate committee signals
- 5:51something profound here. It shows that the traditional
- 5:54tools we have for protecting intellectual property.
- 5:57Like copyrights and patents. Yeah, copyrights,
- 5:59patents or even basic terms of service. They
- 6:02are completely failing in the AI age. So, OK,
- 6:04if software IP is suddenly impossible to defend
- 6:07from distillation like your digital moat is basically
- 6:11evaporating. Is that why these same companies
- 6:14are suddenly pouring billions into proprietary
- 6:16silicon? That's exactly it. Are they basically
- 6:18trying to build a physical moat because the digital
- 6:21one just totally failed? That is the exact strategic
- 6:23pivot the entire industry is making right now.
- 6:25Wow. If you can't easily lock down the software
- 6:28weights of your model, well, the next logical
- 6:30step to maintain your edge and your profit margins
- 6:32is to aggressively control the physical hardware
- 6:35it runs on. And that pivot leads us directly
- 6:37to the second major theme in the Kinsoft dispatch,
- 6:41which is this. Yes. Two massive stories landed
- 6:44on the exact same day and both squarely targeted
- 6:48NVIDIA's absolute dominance over the AI silicon
- 6:51market. It was a big day. Yeah. First up, OpenAI
- 6:54and Broadcom unveiled Jalapeno, which is OpenAI's
- 6:59very first custom in -house AI inference processor.
- 7:03But here's where it gets really interesting.
- 7:05The dispatch reports that this chip went from
- 7:07its initial design phase to tape out. In just
- 7:10nine months. Right. Nine months. And they're
- 7:12planning to deploy it in gigawatt scale data
- 7:14centers by the end of the year. I have to push
- 7:17back on this timeline a bit. Sure. Nine months
- 7:19from design to tape out. That feels like it breaks
- 7:22the laws of physics and semiconductor manufacturing.
- 7:24Yeah. Even, you know, insanely well -funded teams
- 7:26take years to get silicon ready for the foundry.
- 7:29Yeah. Usually they do. How is that even mechanically
- 7:31possible? Well, it sounds impossible if you assume
- 7:34they started from a blank whiteboard. OK. They
- 7:36didn't. To pull off a nine month tape out. OpenAI
- 7:40heavily leveraged Broadcom's existing pre -verified
- 7:44IP blocks. Ah, I see. Think of it like building
- 7:48a house with prefabricated walls and plumbing
- 7:50rather than mixing the concrete yourself on site.
- 7:54Broadcom already has industry -leading SUEs,
- 7:57which are these serializer, deserializer blocks
- 7:59that handle high -speed data transfer between
- 8:02ships. Right, the networking, basically. Exactly,
- 8:05and they already have proven memory controllers.
- 8:07So OpenAI essentially brought the specific matrix
- 8:11math architecture they needed for their AI, and
- 8:14Broadcom... just wrapped it in their existing
- 8:16foundational tech. So they basically snapped
- 8:18these high -end Lego pieces together. Pretty
- 8:21much, yeah. But even then, I mean, why go through
- 8:23the massive headache of building this when they
- 8:25already have access to NVIDIA's top -tier hardware?
- 8:28Like, they already have the best stuff. It really
- 8:30comes down to the architecture of inference versus
- 8:33training. Okay, break that down for me. So...
- 8:36NVIDIA's flagship GPUs are incredible general
- 8:40-purpose AI accelerators. They are built to handle
- 8:43the intense computational precision you need
- 8:46for backpropagation during the training phase.
- 8:48Right, when the model is actually learning. Yes.
- 8:51But Jalapeno is an inference -only chip, meaning
- 8:54when a model is just answering user prompts in
- 8:57production, you don't actually need that massive
- 9:00compute density. What you need at that stage
- 9:03is memory bandwidth. You need to move weights
- 9:05in and out of SRAM as fast as possible with the
- 9:08lowest possible latency. Okay, so it's a completely
- 9:11different job. Exactly. By designing a chip strictly
- 9:14for inference, OpenAI strips away all that silicon
- 9:17real estate dedicated to training and replaces
- 9:20it with pure memory throughput. Which means they
- 9:22run their models vastly cheaper. faster and cooler
- 9:26than using general purpose hardware. Absolutely.
- 9:29Okay. It makes total sense for them to carve
- 9:30out independence on the physical silicon side.
- 9:32But then literally on that exact same day, we
- 9:36get the second half of this rebellion. The software
- 9:38side. Right. Qualcomm confirmed it's acquiring
- 9:41an AI software startup called Modular for over
- 9:44$3 .9 billion. A huge acquisition. Massive. And
- 9:49the dispatch says Modular created a hardware
- 9:52agnostic runtime that attacks NVIDIA's software
- 9:55lock -in. Which we know is CDA. Yes, CDA has
- 10:00been the gold standard. But again, let's get
- 10:02into the mechanics here. People always throw
- 10:04around the phrase universal translator for software
- 10:06like this. But how does a hardware agnostic runtime
- 10:09actually decouple the code from the chip? Well,
- 10:11if we connect this to the bigger picture, we
- 10:13really have to look at the compiler level. For
- 10:15over a decade now, if a developer wrote AI code
- 10:18in a framework like, say, PyTorch, the absolute
- 10:21easiest way to make it run fast was to use NVIDIA's
- 10:25proprietary CUDA libraries because they are highly
- 10:28optimized for NVIDIA's specific GPU architecture.
- 10:31Right. It created a massive walled garden. You
- 10:34basically had to use their hardware if you use
- 10:36their software. Exactly. It's a complete ecosystem
- 10:38lock -in. Modular tackles this by utilizing an
- 10:41intermediate representation layer, specifically
- 10:44something called MLIR, or multi -level intermediate
- 10:48representation. So it basically sits between
- 10:50the high -level Python code and the actual metal
- 10:53of the chip. Exactly that. Instead of compiling
- 10:56directly to CUDA, Modular's engine takes the
- 10:59AI workload, breaks it down into this intermediate
- 11:01mathematical representation, and then uses custom
- 11:04compilers to translate it perfectly into the
- 11:07native instructions. set of an AMD chip, an Intel
- 11:10accelerator, or even, you know, a custom chip
- 11:12like Jalapeno. Wow. It mathematically decouples
- 11:15the software layer entirely from the physical
- 11:17hardware layer. So what does this all actually
- 11:19mean for the listener? I mean, we have this incredible
- 11:21pincer movement happening here. Yeah, a two -pronged
- 11:24attack. Jalapeno attacking NVIDIA's hardware
- 11:26dominance by proving companies can build custom
- 11:29silicon rapidly. And then Modular attacking the
- 11:32software dominance by breaking that CDA lock
- 11:35-in. Mm -hmm. Why should a business leader or
- 11:38just an everyday user care about this highly
- 11:40technical silicon war? Because this fierce competition
- 11:43is the exact mechanism that will drive down the
- 11:46operational costs of the entire digital economy.
- 11:49It always comes back to cost. Always. When developers
- 11:52aren't forced to buy one specific brand of hardware,
- 11:55and when companies like OpenAI aren't paying
- 11:58massive premiums for general purpose inference
- 12:00compute, the unit economics of AI just... plummet.
- 12:05Right. For you, the listener, it means AI integrations
- 12:08in your enterprise software, your customer service
- 12:11tools, your daily workflows. They will all become
- 12:14significantly cheaper and exponentially more
- 12:16ubiquitous. You are literally watching the monopolistic
- 12:20tax being stripped out of the ecosystem in real
- 12:22time. OK, so we are rapidly building out gigawatt
- 12:25scale data centers for these custom chips. Yes.
- 12:27And we're deploying massive new software compilers
- 12:30to route these workloads. Oh. But as we pour
- 12:32the concrete for these data centers and write
- 12:34millions of lines of new orchestration code,
- 12:36the attack surface for bad actors expands massively.
- 12:40Oh, without a doubt. And that brings us to the
- 12:41final and frankly, the most urgent section of
- 12:44the Kinsoft Dispatch today, which is the cybersecurity
- 12:46warnings. It is just the unavoidable physical
- 12:49reality of infrastructure expansion. You cannot
- 12:52scale data centers to the gigawatt level and
- 12:55rewrite your entire software stack without introducing
- 12:58new vulnerabilities. And the vulnerability highlighted
- 13:01in this dispatch is just profoundly ironic. It
- 13:05really is. I mean, we just spent 20 minutes talking
- 13:07about autonomous reasoning models and hyper advanced
- 13:09AI silicon. But on June 25th, the American Cyber
- 13:13Agency, CISA, issued an urgent patching directive
- 13:16for a phone system. Yep. iPhone system. Specifically,
- 13:21Cisco Unified Communications Manager. The bug
- 13:24is cataloged as CVE -2026 -2030. It's an unauthenticated
- 13:30flaw that let attackers write files directly
- 13:32to the underlying system, and it was being actively
- 13:35exploited in the wild. Actively exploited, which
- 13:37is why the June 28th deadline was so hard. Right.
- 13:39But I have to admit some confusion here. Why
- 13:42do sophisticated threat actors, who could theoretically
- 13:45be targeting these massive AI clusters, care
- 13:47about a boring legacy office phone system? Well,
- 13:50this raises an important question about network
- 13:52topology and something called lateral movement.
- 13:54Attackers don't usually try to breach the most
- 13:56heavily guarded fortress door, right? They look
- 13:59for the forgotten side window. The easy way in.
- 14:01Exactly. A unified communications system is essentially
- 14:05a server that sits right on the edge of your
- 14:08network. It has to connect to the public internet
- 14:10to route voice traffic. Sure. But it also has
- 14:13deep, trusted connections to your internal enterprise
- 14:16network to access things like user directories
- 14:19and databases. Ah. So it's basically a bridge
- 14:23between... the wild west of the Internet and
- 14:25the safe internal corporate network. Yes, exactly.
- 14:28And because it's an appliance that just quietly
- 14:30works in the background, IT departments often
- 14:33treat it as a set it and forget it device. Right.
- 14:35It just gets pushed to the bottom of the patching
- 14:37queue. Right. And when an attacker exploits an
- 14:40unauthenticated file write flaw on that specific
- 14:43appliance, they aren't just messing with your
- 14:45voicemails. They are dropping a web shell. Oh,
- 14:47wow. They gain root access to that server. And
- 14:49from there, they pivot. They use that trusted
- 14:52appliance to move. laterally inside the network,
- 14:54targeting the domain controllers, deploying ransomware
- 14:57or, you know, exfiltrating data. Yikes. So the
- 15:01grounded practical takeaway for you listening
- 15:03is this. You absolutely must treat every single
- 15:06piece of internet -facing enterprise software,
- 15:09especially the mundane legacy appliances like
- 15:13a phone system, as a critical infrastructure
- 15:16vulnerability. Which honestly feels a bit overwhelming.
- 15:19It is. I mean, the sheer volume of code in modern
- 15:21enterprises makes manual patching feel like trying
- 15:24to empty the ocean with a teaspoon. Exactly.
- 15:26But the dispatch does introduce a massive shift
- 15:30in defensive capabilities to counter this, right?
- 15:32It does, yeah. Because that same week. OpenAI
- 15:35expanded its Daybreak cybersecurity program.
- 15:38They released a specialized security -focused
- 15:40AI model to vetted defenders. And the pitch here
- 15:44isn't just an AI that flags anomalies, right?
- 15:46No, it goes much further than that. It includes
- 15:48tools for automatically finding and patching
- 15:50vulnerabilities in code at scale. So instead
- 15:53of a human security analyst endlessly hunting
- 15:55for zero days, the AI does static analysis, understands
- 15:59the semantic context of the code, and basically
- 16:01writes the patch itself. And this brings us to
- 16:04the terrifying duality of autonomous agents and
- 16:06cybersecurity. Right. Because the concept of
- 16:09automated patching, where a system discovers
- 16:11a buffer overflow and writes the mitigation code
- 16:14in literally milliseconds, it's structurally
- 16:17necessary for defenders to keep up at this point.
- 16:20Yeah, because humans are just too slow. Right.
- 16:22The time to patch window is basically shrinking
- 16:25to zero. Yeah. But the dispatch cautions heavily
- 16:28about the dual use nature of this technology.
- 16:30Right. Because an AI model that is sophisticated
- 16:32enough to read millions of lines of code, find
- 16:35a deeply hidden logic flaw and understand how
- 16:38to patch it. Yes. Is also, by definition, sophisticated
- 16:41enough to exploit it. Exactly. We are transitioning
- 16:45right now into a machine speed arms race. The
- 16:48same static analysis capabilities that empower
- 16:51the Daybreak program to secure an enterprise
- 16:53can be entirely inverted by threat actors. Wow.
- 16:56They can use it to discover zero -day vulnerabilities
- 16:58in proprietary software before the vendor even
- 17:01knows they exist. That is just wild to think
- 17:04about. We are moving away from the era of human
- 17:07operators manually probing networks into this
- 17:11era where automated AI tools are simultaneously
- 17:13defending and attacking infrastructure at speeds
- 17:17human analysts simply cannot process. Wow. OK,
- 17:21let's take a breath and just bring all of this
- 17:23together for you. We started this deep dive looking
- 17:26at a single dispatch from June 2026. But it's
- 17:30not just a laundry of tech news. It is a high
- 17:33definition snapshot. Absolutely. Right. Yep.
- 18:00And finally, we looked at the massive security
- 18:01fallout of this expansion, balancing the urgent
- 18:04reality of lateral network attacks via Cisco
- 18:06phone systems against the futuristic machine
- 18:09speed promise of OpenAI's Daybreak program. It's
- 18:12a lot to take in. And as we consider the architecture
- 18:15of this new era, I want to leave you with a final
- 18:18thought to mull over, building on that last point
- 18:20about security. Okay. We discussed the duality
- 18:23of AI and cybersecurity and that shrinking time
- 18:26-to -patch window. Well, if AI systems can automatically
- 18:29find and patch vulnerabilities at scale, as Daybreak
- 18:32promises, what happens when two opposing AIs,
- 18:35one designed to relentlessly attack and one designed
- 18:38to autonomously defend and engage in an automated
- 18:41cyber war at speeds no more than a second, No
- 18:42human can monitor. Oh, wow. Will our enterprise
- 18:45networks and the basic Internet infrastructure
- 18:47we rely on every day simply become passive battlefields
- 18:51for autonomous agents fighting invisible millisecond
- 18:53wars? That is a wild and honestly somewhat chilling
- 18:56thought to leave on. The landscape is entirely
- 18:59new and the fundamental infrastructure of the
- 19:02digital world is being rewritten as we speak.
- 19:04We hope this deep dive gave you the mechanistic
- 19:07context you need to truly understand these shifts,
- 19:09because taking the time to explore the how and
- 19:11why is exactly how you stay ahead of the curve.
- 19:14Thanks for joining us today. Stay curious and
- 19:16stay patched.