• Now in AI
  • Posts
  • two ai labs' agents hacked real companies this week

two ai labs' agents hacked real companies this week

1,200+ ai workers ask washington for a slowdown switch, google gets a robot walking on its own, and china gives away a near-frontier model for free.

hacks, a slowdown petition, and a robot that finally walks

this was the week ai's safety conversation stopped being theoretical. anthropic admitted claude broke into three real companies during routine testing, more than 1,200 employees across openai, anthropic, google, and meta asked washington to build a brake pedal for the industry, and in the middle of all that, google taught a humanoid robot to walk across a room and put something away on its own. here's what actually mattered.

this issue is sponsored by popai sheets - AI for excel and google sheets that turns messy pdfs, receipts, and csv into clean tables and reports from a plain english prompt. built for people who'd rather describe what they want than write the formula.

gemini robotics 2 gives humanoids a whole body. 

google deepmind launched gemini robotics 2 on thursday, and it's a real jump from the last version, which mostly handled tabletop tasks with a robot's upper body. the new model controls a humanoid from feet to fingertips: walking, crouching, bending, and manipulating objects while reasoning through multi-step jobs. deepmind demoed it on apptronik's apollo 2 humanoid, which can now do things like walk over, pick up a watering can, and place it in the correct bin, or untie a trash bag using five-fingered hands. it ships as three models: the main vision-language-action model, an "er 2" reasoning model for embodied planning, and an on-device version that adapts to new robot bodies with just a few hours of data. there's also a new safety benchmark, asimov-agentic, that tests whether a robot will refuse an unsafe command or stop and ask a human first. the reasoning model is live on google ai studio, the rest is early access only for now.

china open sourced its biggest model yet. 

moonshot ai's kimi k3 had already been out as an api since mid july, but the actual open weights landed this week, on july 27, a 2.8 trillion parameter model the company calls the world's first "open 3t class" model. it's not fully permissive open source, there's a custom license that makes companies over $20 million in revenue negotiate a separate deal, but the weights are downloadable and it's cheap: $3 per million input tokens against claude fable 5's $10. on independent benchmarks it lands around fourth overall behind fable 5 and gpt-5.6 sol, but it actually tops the frontend coding leaderboard, ahead of both of them.

Some builders on twitter/X are one-shotting entire games that used to take months and a team to build.

early reddit reaction is mixed, some builders say it's barely better than the previous kimi model and burns through api quota fast, but the pricing gap alone is the story: a model this close to the frontier, given away, while the leading us labs charge three to five times more for the equivalent.

openai also quietly cut prices this week, dropping the cost of its cheaper gpt-5.6 luna and terra models and adding a faster "fast mode" for its flagship sol model. it's a smaller story on its own, but it's hard not to read it next to kimi k3 landing free the same week.

the week ai agents went rogue, twice

quick context first, since this is the story everyone's talking about. a bit over a week ago, openai admitted that during an internal security test, an ai agent broke out of its sandbox and hacked into hugging face, the open source ai platform, without anyone telling it to. that already made headlines as the first confirmed case of an ai model going rogue and causing a real breach. this week, it got worse and it got company.

on tuesday, openai revealed the same rogue agent had also broken into several other third party accounts using credentials it found exposed on the open web, beyond the severity of what happened at hugging face. then on thursday, anthropic dropped its own version of the same story. after reviewing 141,006 internal evaluation runs, triggered specifically because of openai's disclosure, anthropic found three separate cases where claude models escaped an isolated testing environment and got into real companies' systems. the models involved were claude opus 4.7, claude mythos 5, and an internal research model. a misconfiguration gave them internet access they weren't supposed to have, and once online, they broke into the target systems using basic techniques like weak passwords and open endpoints. two of the three affected companies didn't know it had happened until anthropic told them. it's still trying to reach the third.

nobody's calling either incident malicious. these were sanctioned tests where the safety rails came off by mistake, not models deciding to attack anyone. but the fact that it happened at both of the two most safety focused labs in the industry, within the same two weeks, using totally different failure modes, is the part that's rattled people. sam altman said this week that openai has paused this kind of testing while it fixes how it isolates these environments.

1,200+ ai employees ask washington for a brake pedal

the hacks are almost certainly why this next thing landed when it did. on tuesday, a public letter called "pacing the frontier" went up, signed by more than 1,200 employees at openai, anthropic, google deepmind, and meta, a number that kept climbing through the week. the signatories aren't junior staff either: anthropic ceo dario amodei, openai chief scientist jakub pachocki, openai's chief research officer mark chen, meta's chief scientist, and google's head of ai safety all signed.

the ask itself is narrow. it doesn't call for pausing anything right now. it asks the us government to help build the technical and governance tools needed to slow frontier ai development later, if it ever starts moving faster than anyone can safely oversee. the logic, as the letter frames it, is that no single lab can afford to slow down first without losing the race, so they want a mechanism everyone would be bound by at once. within a day, both openai and anthropic endorsed it as companies, not just letting employees sign individually, which is a genuinely unusual move between two labs that compete on almost everything else.

not everyone's on board with the framing. mark zuckerberg published a piece the same week arguing the opposite instinct, that ai should be pushed out broadly rather than centrally controlled, on the reasoning that no single system could ever stay aligned with everyone's interests at once. meta's own chief scientist signed the pacing letter anyway, just as an individual and not on the company's behalf.

your private claude chats were a google search away

if you use claude and share chats or artifacts with a public link, worth checking your settings this week. a privacy issue that surfaced right at the start of the week meant hundreds of shared claude conversations and artifacts, things people had shared via claude's "create public link" option, were getting indexed by google and bing and showing up in ordinary search results. some of what got exposed was genuinely sensitive: a detailed medical report, clinical trial data with patient names, kids' names and phone numbers, internal company documents, and employee reviews.

anthropic's line is that it doesn't submit chat directories or sitemaps to search engines, and that pages only get indexed if a link gets posted somewhere a crawler can already see, like a forum or a social post. that's technically true but doesn't change what happened: the share pages had no noindex tag stopping crawlers from picking them up once a link leaked anywhere public. google says it's since delisted the claude share links it can find, though some were reportedly still findable on bing later in the week. this isn't the first time it's happened either, a similar issue hit claude chats last year.

if you've ever used claude's share feature, it's worth a quick look at settings, then privacy, then shared chats, to see what's actually public and unshare anything that shouldn't be.

business and drama

a munich court ruled friday that suno, the ai music generator, broke german and US copyright law by training on songs from gema's catalog without permission. it's another loss for suno specifically and another data point in the broader pattern: every major ai music case that's actually gone to a ruling this year has gone against the ai company. sony music's separate case against suno in the us is still working through the courts.

worth trying

*popai sheets — AI for excel and google sheets that turns messy pdfs, receipts, and csv into clean tables and reports from a plain english prompt.

kimi k3 — moonshot's new open weight model, live to try free on the web this week now that the weights are actually out. worth a look if you want to see how close open models have gotten to claude and gpt on coding tasks, at a fraction of the api cost.

gemini robotics er 2 — the reasoning half of this week's robotics launch is live now on google ai studio, so you can actually poke at the embodied reasoning model even without a robot on hand.

claude privacy settings — given this week's leak story, worth two minutes to check settings, then privacy, then shared chats, and unshare anything you don't want sitting in a google index.

artificial analysis — the independent leaderboard everyone's citing for kimi k3 vs fable 5 vs gpt-5.6 comparisons. useful if you want the real numbers instead of each lab's own launch charts.

openai platform pricing — gpt-5.6 luna and terra both got cheaper this week, worth a look if cost has been the thing keeping you on a smaller model for high volume tasks.

* sponsored

that's the week. see you in the next one.