
Нам всем уже надоели новости об очередных «достижениях» LLM. Но, несмотря на поток надоедливых новостей, LLM представляют собой одно из ключевых направлений современных технологий (пусть и несколько переоцененное), и здесь иногда происходят некоторые открытия, которые становятся значимыми вехами...
Я изучаю историю технологий и не мог пройти мимо статьи, которая говорит об обнаружении внутри мощных LLM паттернов-концепций, которые активируются, когда модель «думает» о связанных с этим паттерном вещах. И, что самое интересное, модель может «осознать» этот паттерн. А исследователи могут посмотреть в эту область и понять, пытается ли модель жульничать и что у нее реально на уме...
Считаю, что русскоязычный читатель должен иметь возможность ознакомиться с этим текстом. Это полный перевод на русский язык свежей, от 6 июля 2026г. статьи на сайте Anthropic
Читать далее
OpenAI was evaluating its artificial intelligence models’ ability to exploit vulnerable software when instead the models hacked the infrastructure surrounding the test, broke containment, and attacked a real company, OpenAI revealed on July 21. Observers say this is the first real-world instance of AI doing something researchers have long worried about: a loss-of-control scenario. If the industry fails to learn from it, it is unlikely to be the last.
Hugging Face, a company that hosts AI models and datasets, was the target of the autonomous attack, and reported the incident to local police before it knew OpenAI’s models were responsible. The breach was serious, but the immediate consequences were limited. Had similar behavior occurred inside a hospital, power grid, or other critical system, it could have been much worse.
Many have called the incident a “warning shot.” TIME spoke with experts and insiders about what it would take to heed that warning before a similar failure produces consequences that are harder to contain.
On July 16, Hugging Face said it had been hit by an unusually automated cyberattack. Over the course of a weekend, AI agents carried out thousands of actions across many temporary virtual computers, moving through the company’s internal systems and shifting the infrastructure coordinating the attack between online services to keep it running. Five days later, OpenAI disclosed that its own models were responsible.
The models were trying to cheat on a cybersecurity test. OpenAI had placed them inside what it called a “highly isolated environment,” with only limited access to an internal service used to download approved software. They found a previously unknown flaw in that service, used it to break into other OpenAI systems and eventually reached the open internet. From there, they inferred that Hugging Face might hold material related to the test, broke into its systems and obtained information that helped them score higher.
“If a model of this capability level cannot be contained, what should we expect for future, much more powerful models? This is an important wake-up call both for risks from loss of control of powerful AI systems as well as organizational security for frontier labs,” says Marius Hobbhahn, CEO and founder of Apollo Research, which tests AI models for deception and scheming.
How long were the agents running? Did they work in unison? What was the prompt? These details remain unknown, at least to the public. Several experts TIME spoke with stressed that the lack of details make the severity of the incident hard to judge.
OpenAI did not respond to TIME’s request for comment, but has said it has partnered with Hugging Face to conduct a thorough investigation and will share more details once complete. But perhaps more worryingly, OpenAI is not legally compelled to disclose the incident in the first place.
California’s SB 53 and New York’s RAISE Act are recently passed state-level laws that require large AI companies to disclose critical safety incidents, but only if an incident risks causing more than 50 deaths or serious injuries, or more than $1 billion in property damage. “They have made the bar so high for anything to qualify, only the most grievous incidents will actually be reported,” says Mackenzie Arnold, director of U.S. policy at LawAI, an independent think tank focused on the legal challenges posed by artificial intelligence.
“The version of the RAISE Act that the NY Legislature passed would have required disclosure of this ‘incident.’ After lobbying from OpenAI, Bloomberg, and a16z, the final version the Governor signed allows companies to hide events like this,” Alex Bores, New York state representative and the bill’s sponsor posted to X. “I'm glad OpenAI chose to disclose this crime. The law shouldn't give them a choice,” he wrote.
If AI can slip out of its box, you need a stronger box.
“Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,” says an OpenAI staffer, who spoke under the condition of anonymity. The day before OpenAI disclosed the incident, the company revealed that it had shut down another internal deployment after it realized it had slipped out of its sandbox—a digitally, rather than physically, separated environment. “Models have broken out of sandboxes before, and we always try to patch them,” the staffer says. “But the problem is … it's impossible to patch every single thing that a creative AI can do.”
The problem is not unique to OpenAI. Anthropic disclosed in April that it realized an internal deployment of Mythos had gained unauthorized access after one of its researchers received an email from the model while having lunch in a park.
As models surpass human ability to craft escape routes, the challenge of containment becomes daunting. “I think we do need to do that, but I think we also need to be prepared for a world where even best practices aren't really good enough,” says Peter Wildeford, head of policy at the AI Policy Network, a nonprofit that advocates for policies to help America prepare for a world where AI matches human cognition.
But while in theory, a superhuman AI system might outmaneuver even the most secure containers, there are also actions AI companies could take right now to improve security. Currently, there are no laws governing the security of internal deployments.
“Sandboxes are actually notoriously insecure,” says Heidy Khlaaf, chief AI scientist at AI Now Institute, and a former safety systems engineer contractor at OpenAI. The fact that the models were permitted to connect to a service for downloading packages meant the environment was not truly sealed off, she adds.
In a previous role, Khlaaf audited security at dozens of technology companies. Before that, she worked auditing high-risk systems, like those used inside nuclear power plants—which often “air gap” systems, physically cutting them from internet access. “What we consider safe in a nuclear plant is so different from what big tech considers safe.”
The Hugging Face incident reveals the importance of real-time monitoring.
Though details of the precise timeline are scant, Hugging Face has said the agents worked over a “weekend,” suggesting that they were able to break containment and get up to no good for an extended period before OpenAI noticed and intervened. Actions carried out internally by agents on OpenAI’s Codex platform are carefully monitored, the OpenAI staffer says, but models undergoing evaluation are deployed on a separate system that is not monitored by default.
Zack Korman, CEO of Oslo-based agent-oversight startup Embroidery, says real-time agent monitoring—even outside top AI companies—is commonplace, and to not carefully oversee a cybersecurity evaluation is “irresponsible.” You should be confident models cannot break free, “but also have monitoring just in case you're wrong,” he says.
OpenAI has said “this incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing.”
To prevent models from taking actions, OpenAI typically installs guardrails on its models after training to reduce the chance they’ll engage in harmful actions. In this instance, those cyber guardrails were disabled to properly measure its performance.
But the ultimate aim of the field of “AI alignment” is to ensure that such guardrails become less necessary as AI models naturally behave as intended. This is seen as particularly important to those who believe AI may become smart enough to route around guardrails.
“We train the models to be really good at accomplishing tasks and doing whatever it takes to accomplish those tasks,” the OpenAI staffer says. What remains an open technical question is how to guarantee those models don’t take unintentional or dangerous actions. “We're still nowhere near solving this misalignment problem,” they add.
There’s another way the industry could steer development in a safer direction, Khlaaf says. While some abilities, like spotting vulnerabilities in software, are useful to attackers and defenders, designing exploits for those vulnerabilities uniquely empowers attackers. Khlaaf says that labs should put more resources into capabilities that directly help defenders, such as detecting attacks, writing secure code, and patching vulnerabilities. AI companies are already investing in some of those tasks, but their most visible capability gains and benchmarks have centered on finding vulnerabilities and constructing exploits, she says.
In an industry defined by speed, OpenAI has said the stricter infrastructure controls it has implemented in response have already slowed its “research velocity.” Hobbhahn says that’s a price worth paying. “This is humanity’s last technology. We cannot screw this up. So we need to err on the side of getting it right rather than getting it immediately.”

Рано или поздно к вам приходит сотрудник с запросом о повышении зарплаты. Вы видите: за прошедший период он многое сделал, заметно прокачал технические и организаторские навыки. Сомнений в том, что он может пойти на повышение грейда, нет. Но прежняя шкала оценки уже не годится: она не учитывает умение работать с новыми инструментами. ИИ изменил разработку, однако осознали это далеко не все.
В этой статье я расскажу о том, как я вижу современную разработку. Обсудим, что важно понять нам самим и что донести до сотрудника, чтобы минимизировать риск его ухода и сохранить мотивацию.
Читать далееOver 11,500 people this summer are suspected to have been sickened by disease that can cause explosive diarrhea, CDC says
Investigations into the largest cyclosporiasis outbreak on record expanded to nine states this week, as more than 11,500 people are suspected to have been sickened this summer, according to reports and the Centers for Disease Control and Prevention (CDC).
Data released by the CDC this week said the agency is looking into 4,173 laboratory-confirmed domestic cases of the disease and investigating an additional 7,400.
Continue reading...Since I saw the movie, I have found it difficult not to see Trojan horses everywhere. We live now in a state of perpetual distrustfulness
Prior to the release of Christopher Nolan’s The Odyssey, I’m not sure many of us would have conceived of the Trojan Horse as a kind of proto-atomic bomb, the sort of norm-obliterating device from which civilisation cannot turn back. If anything, the horse has long existed in our collective imaginations as at best a credulity-stretching plot device – why did those overly trusting Trojans not think to look that particular horse in its wooden mouth? – and at worst a cutely comedic non-threat. The greatest civilisation of its age was taken down by a timber what?
But the true masterstroke of Nolan’s film is its parallels to the last one he dropped on us, Oppenheimer. Like that film, The Odyssey is about a haunted genius who, when presented with a problem, solves it with such speed and significance that he must necessarily cast aside implication. But rather than splitting the atom, Odysseus (Matt Damon), dismembers something considerably more significant – the social fabric.
Continue reading...Streaming pioneer that transformed TV faces slowing growth and battle for viewers’ attention from YouTube
When Netflix offered viewers the chance to binge on an entire season of House of Cards it revolutionised the TV industry and started on a path to becoming the world’s most popular streaming service.
Almost 15 years on, however, the upstart’s strategy of pumping its service with an avalanche of content and an almost-as-rapid penchant for mercilessly cancelling shows has viewers reaching for the remote.
Continue reading...
This spring, I was among the first journalists to speak with Christopher Nolan about his highly anticipated adaptation of The Odyssey. Now, the movie is out and has become Nolan’s biggest worldwide opening ever.
In our initial interview, Nolan was reluctant to discuss the ending of The Odyssey. He’s known as the master of the last-act twist, and likes leaving the audience with more questions than answers. He told me he never wants to impose his interpretation on others. (Are Cobbs’ kids in Inception real or imagined? You decide.) But when I hopped on the phone again shortly before the film’s release to have a spoiler-filled chat about the film’s final act—readers beware!—he offered some insight as to why he made certain changes when adapting the epic poem.
We discussed the surprising parallels between The Odyssey and Oppenheimer, why Nolan decided to give Odysseus and Penelope the happy ending they don’t get in Homer’s epic poem, and how to interpret Zendaya’s mysterious character: Is she actually a goddess or merely a manifestation of Odysseus’ guilt? We also dug into Zeus’ Law and the deeper meaning behind Helen’s scarred face.
Nolan: My form of storytelling is, you take a hero, as we did with Oppenheimer, and Oppenheimer has all these flaws, all of these complexities. You find an actor who can take the audience on a journey with that character and make mistakes with that character and not judge that character. And then you step back from that character and go, “Well, how did we get here?”
[For this movie,] you do that through Odysseus’ men. That’s one of the things in the poem that’s breezed over—the way in which the crew ends up getting killed one by one, despite his leadership. In the process of adaptation, you go, “That’s actually pretty interesting. Maybe we need to make that more center stage.” What’s the relationship between him and his crew? How do they feel about how he’s leading them and the things that befall them?
He’s a guy who makes a lot of mistakes. His ultimate mistake is the horse. So there’s a sort of tragic flaw right from the beginning in our telling.

I couldn’t possibly answer that because I learned very early on with Memento that it’s unfair to define the experience for the audience by answering questions like that.
Here’s what I will say: For me, what was important was that we view the gods the way that people at the time might have viewed them. So they’re seeing evidence of the gods in nature. They’re seeking and finding gods and the influence of gods in the people around them. Odysseus says to Telemachus, “Don’t look for gods in men. You’ll just be disappointed.” And some of what Odysseus says about divinity, to me, is somebody who’s protesting a bit too much, someone trying to live in denial of their own instincts and their own practices.
Yes. If I’ve done a film right, it leaves me with questions. It leaves me with things I want to keep exploring and thinking about. And Oppenheimer fed very much into my reading of the story of Odysseus.
John Leguizamo [who plays Odysseus’ friend Eumaeus] says it at the very beginning: “That’s very clever, but your cleverness will get you in trouble.” As it does again and again and again. This is where it ties into Oppenheimer: Just because you can doesn’t mean you should. The parallels were pretty apparent to me as I figured out how to adapt it.
When you find these things are difficult to deal with, sometimes the solution is just to dive in and make it everything. I didn’t use the term xenia because you don’t want to throw too many terms at people. But it’s essentially the golden rule. It all comes back to weather, to geography. When you leave the house, you are depending on the kindness of strangers, and you have to trust that when you arrive somewhere after weeks or months of traveling, that you will be fed, you will be given water. Three thousand years ago, that was everything. You were always at the mercy of strangers. And so this idea that a stranger could be a god in disguise, it’s very powerful as the underpinning of a belief system.

That line, “The face that launched 1,000 ships, or maybe just 500” just occurred to me to write. Sometimes what you do as a writer is purely instinctive. I felt the poem didn’t necessarily address what it is for Helen to be brought back by Menelaus and how that would ever work. How would that relationship be? How could it possibly be anything other than the most acidic, awful marriage?
There’s a lot of sh-t that’s happened in that relationship, and we have a very brief moment of time with them in this story. I was looking for a shorthand way to communicate the aftermath of the war and how it’s been traumatic for the entire world.
Circe articulates it well in the film. I love the way Samantha [Morton] has played the part. Talking about choices of adaptation, the end of the poem is so grim and so misogynistic, I had to exclude some of that material. But I didn’t want to completely take it off the table in terms of what it is saying about men and the way they wage war.
Well, it is the fall of civilization, so there is a fairly bleak sensibility. The end is talking about cycles, and you can view cycles pessimistically: History is doomed to repeat itself. But there is also something optimistic about those cycles. As Penelope says, “Civilization will rise again.” There is rebirth and regrowth.
For me, it was important in looking at the ending to try and strike the right balance between pessimism and optimism—pessimism about what might have precipitated the Bronze Age catastrophe and the optimism that positive things will reassert themselves.

Я включил на домашнем роутере файловый журнал DNS-запросов и оставил на выходные. За 39 часов набралось 106 875 обращений к 2605 доменам от восьми устройств: компьютера, ноутбука, телефона, умной колонки, робота-пылесоса, телевизора и стиральной машины.
Дальше начались открытия.
Умная колонка выходит в сеть раз в 43 секунды и не останавливается никогда — ночью столько же, сколько днём. Причём обращений в аналитику у неё ровно столько же, сколько к рабочему API.
Робот-пылесос обошёлся тремя доменами за двое суток и не позвал никого постороннего, хотя ни разу не убирался.
Стиральная машина ходит только в облако производителя — четыре домена. А приложение для управления ею принесло с собой три сторонних аналитических сервиса.
Телевизор оказался сложнее всех. Ночью он спит крепче остальных приборов, а за сутки простоя успевает обратиться к 172 доменам. Среди них — компания, которая занимается распознаванием картинки на экране. Причём это подтверждается не догадками, а политикой конфиденциальности самого производителя телевизора.
В статье: полный разбор по каждому устройству с цифрами, две грабли при настройке журнала на роутере с 44 МБ свободного места, и все команды, чтобы повторить у себя за вечер.
Читать далееThe One Nation leader has allied herself with men’s rights activists for years, consistently suggesting women routinely make false allegations of violence
Get our breaking news email, free app or daily news podcast
For a decade, Pauline Hanson has been beating the same drum.
Her comments this week – suggesting domestic violence was a “two-way street” and questioning why women did not simply leave abusive relationships – have caused significant concern and been labelled “dangerous” by experts and other commentators.
Continue reading...Economists’ warning comes as escalating Middle East crisis pushes global crude oil back above $US100 a barrel
Get our breaking news email, free app or daily news podcast
Australian households face the prospect of a Reserve Bank interest rate hike and petrol prices above $2 a litre over the coming weeks, economists warn, as the escalating Middle East crisis pushes global crude back above $US100 a barrel.
Financial markets now see an even chance the RBA board will deliver a fourth cash rate increase at the next meeting on 11 August.
Continue reading...Человек, чей SeaBIOS бутил ваши виртуалки с 2010 года, сегодня единолично решает судьбу прошивки ваших 3D-принтеров. У Klipper один мейнтейнер, один открытый issue — табличка «трекер закрыт» — и сотни висящих PR. Snapmaker переписал 20% кодовой базы, потому что занести тулченджер в апстрим просто некуда. Разбираю на фактах, как устроен governance де-факто стандарта прошивок для FDM — и во что его отсутствие обходится вендорам, контрибьюторам и вам.
Читать далееExplosive diarrhea infection detected in 41 states with at least 72 sickened by new, as yet unknown source of parasite
Confusion abounds in the cyclospora outbreak that has sickened thousands in at least 41 states, as officials identify a new cluster of cases and more potential products come under scrutiny.
The recall of products during this outbreak of explosive diarrhea has been an unusual one, complicated by poor communication, politicization and the delay of a new regulatory rule could have helped trace the cyclosporiasis outbreak faster.
Continue reading...Over 11,500 people this summer are suspected to have been sickened by disease that can cause explosive diarrhea, CDC says
Investigations into the largest cyclosporiasis outbreak on record expanded to nine states this week, as more than 11,500 people are suspected to have been sickened this summer, according to reports and the Centers for Disease Control and Prevention (CDC).
Data released by the CDC this week said the agency is looking into 4,173 laboratory-confirmed domestic cases of the disease and investigating an additional 7,400.
Continue reading...The conditions for ferocious fires are expected to worsen as the planet heats up, putting more people at risk
Fierce wildfires after back-to-back heatwaves have forced tens of thousands of people across southern Europe to flee their homes, while powerful infernos across Canada and the US choke big cities with thick smoke.
Continue reading...
Google опубликовала выпуск AI & Economy ATLAS — исследование о том, как люди используют ИИ в работе и повседневной жизни. ИИ уже присутствует в широком спектре профессий, но чаще помогает с отдельными задачами, чем заменяет их целиком. Разбираемся, почему так и разбираем особенно интересные моменты отчета.
Читать далееЧеловек, чей SeaBIOS бутил ваши виртуалки с 2010 года, сегодня единолично решает судьбу прошивки ваших 3D-принтеров. У Klipper один мейнтейнер, один открытый issue — табличка «трекер закрыт» — и сотни висящих PR. Snapmaker переписал 20% кодовой базы, потому что занести тулченджер в апстрим просто некуда. Разбираю на фактах, как устроен governance де-факто стандарта прошивок для FDM — и во что его отсутствие обходится вендорам, контрибьюторам и вам.
Читать далееMaresca holds first press conference as City manager
‘It’s a challenge to try to do the right things immediately’
Enzo Maresca said following Pep Guardiola as Manchester City manager is a “privilege” but admitted he is aware how the successors to Sir Alex Ferguson and Arsène Wenger struggled at Manchester United and Arsenal.
Maresca took over at City last month and, at his first press conference in the role, said he understands the challenge of replacing Guardiola, who claimed 17 major trophies in his decade in charge to become the club’s greatest ever manager.
Continue reading...From plot twists to stunning visuals, our flagship long-form journalism can be as gripping as any beach read. To mark the publication of the summer issue of the Guardian Long Read magazine, we look at the craft that goes into making the stories so compelling
For the latest issue of the Guardian Long Read magazine, the brief for the cover was simple: something that evokes the pleasure of summer reading.
“With these magazines, there’s always a front-and-back reveal,” explains Chris Clarke, who has art directed all four issues of the magazine, which collates some of our finest long-form journalism. “For this summer issue, we lean into the joy of undistracted reading: on the front, a woman sits in a boat, reading, as the tide comes in. On the back, she is still there, still reading, except the tide has gone out and the boat is beached on the sand.”
Continue reading...