RSS Feeds

Farage's by-election victory won't stop questions about finances
Published: 2026-08-14 08:18:02 | Created: 2026-08-14 05:33:55
Nigel Farage triggered the Clacton by-election and the result was predictable - so what did he achieve?
show more
Чтение на выходные: книга про успех Pop Mart
Feed: Все публикации подряд на Хабре (https://habr.com/ru/rss/articles/)
Published: 2026-08-14 05:30:29 | Created: 2026-08-14 05:31:55

Сегодня в рубрике — книга о том, как можно создать плюшевое золото. Да такое, что о нём, словно о поп-звезде, услышит весь мир.

Даже если вы не знаете, как выглядит Лабубу, Molly или Skullpanda, информация о ком-то из них наверняка попадала в зону вашей слышимости. А может, приближалась настолько близко, что приходилось раскошеливаться — добровольно или не очень.

У основания этой империи — Ван Нин, выпускник неприметного вуза, чья идея оказалась сильнее денег и скептицизма инвесторов.

Всё началось с отказа. К 2015 году Pop Mart продавала японские фигурки Sonny Angel — они приносили треть всей выручки. Но Pop Mart была лишь дистрибьютором. Права принадлежали японской Dreams. Когда Ван Нин попытался открыть магазин в Шанхае, партнёры запретили продажу. Предложения об эксклюзивной дистрибуции, совместных коллекциях, участии в выставке — всё было отклонено. «Дайте Sonny Angel идти своим путём».Тогда у компании не осталось ничего, кроме как искать особенный путь для Pop Mart.

«Представьте: мы платим огромную аренду, а нам запрещают продавать товар, который дает 30 % всей выручки. Это был удар. Мы пытались договориться: предлагали стать крупным дистрибьютором, заняться маркетингом или запустить совместное производство. Но партнеры были категорически против. Мы столкнулись с полным непониманием», — прокомментировал этот период Ван Нин.

Читателям с нишевыми увлечениями вроде историй брендов и товарных знаков книга будет интересна историями о том, как защищать интеллектуальную собственность. Весь раздел про отношения с японскими партнёрами и поиск своего IP — прости господи, чистый мастрид для корпюристов и стратегов.

Читать далее
show more
'Unprecedented' rain in Japan kills eight people
Published: 2026-08-14 10:36:14 | Created: 2026-08-14 05:29:55
The storm cut power to more than 20,000 households and left 7,000 people stranded at Tokyo's Narita airport.
show more
Farage Wins Special U.K. Election That He Initiated, as Expected
Published: 2026-08-14 11:56:55 | Created: 2026-08-14 05:27:55
Nigel Farage, the leader of Reform U.K., won the Clacton by-election with 63 percent. Count Binface, a novelty candidate, came second with 27 percent, a personal best.
show more
How Maresca's Man City are shaping up compared with Guardiola side
Published: 2026-08-14 05:16:06 | Created: 2026-08-14 05:24:56
Enzo Maresca has the difficult task of following Pep Guardiola at Manchester City. BBC Sport analyses the new manager's style and tactics from the pre-season games so far.
show more
Builders, bakers and cheeseburger-makers: the UK workers struggling in extreme heat
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 05:00:19 | Created: 2026-08-14 05:19:56

From chefs to paramedics, workers tell the Guardian how they are coping and unions outline the changes needed in the workplace

In the sweltering kitchen of the Plimsoll pub in north London, the head chef, Joshua Swinney, has a frozen tea towel wrapped around his neck to combat the heat.

As temperatures exceeded 35C (95F) in the capital, workers who were not lucky enough to be in air-conditioned offices sweltered in the heat.

Continue reading...
show more
Hot topic: why is Burnham so reluctant to talk about the climate crisis?
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 05:00:20 | Created: 2026-08-14 05:19:56

Extreme heat is harming the wellbeing of the ordinary people the PM champions, yet he has ignored the subject

From the first day Andy Burnham came into office, climate scientists, campaigners and some of the ordinary people surviving through this summer of extreme heat began to wonder when he would start to talk about the climate crisis.

Up until this week, many held back from criticism. After all, he had only been in power for a couple weeks; it felt far too early to start to judge and call him out on it.

Continue reading...
show more
Searches for UK homes with air conditioning more than double
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 05:00:20 | Created: 2026-08-14 05:19:56

Most owners would consider installing aircon as heatwaves lead buyers to focus on ability to stay cool, Rightmove says

Searches for homes for sale in the UK with air conditioning have more than doubled this summer as heatwaves push staying cool up the list of priorities for homebuyers, alongside nearness to good schools and transport links.

In the week of the UK’s fifth heatwave, in which temperatures in parts of England hit 38C, the property website Rightmove said the heat was influencing how Britons thought about their homes, with more people paying attention to how comfortable a property would be during hotter weather.

Continue reading...
show more
Nigel Farage wins Clacton byelection in contest boycotted by every other major party
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 05:07:02 | Created: 2026-08-14 05:19:56

Reform UK leader does not appear for the result, despite beating nearest challenger, Count Binface, in byelection he caused amid deepening scandal over £5m gift

Nigel Farage has seen off Count Binface to win the Clacton byelection in a contest boycotted by every other major political party.

The Reform UK leader in July triggered the byelection in an attempt to shake off a deepening scandal over a £5m gift he received from the cryptocurrency billionaire Christopher Harborne.

Continue reading...
show more
NHS service admits data breach due to pager use
Published: 2026-08-14 05:12:44 | Created: 2026-08-14 05:17:55
The medical data of transplant patients from across the UK was sent over an unencrypted pager network.
show more
Son of 102-year-old says father died after being 'pushed off stage'
Published: 2026-08-13 18:28:17 | Created: 2026-08-14 05:16:55
Phillip Ormerod was taken to hospital in critical condition and later died, police have confirmed.
show more
How a major oil slick started washing up on Iran's beaches
Published: 2026-08-14 05:01:00 | Created: 2026-08-14 05:11:55
Experts tell BBC Verify pollution has spread from a ship attacked on 3 August and threatens vulnerable wildlife.
show more
BBC Verify speaks to woman who filmed ICE agent pointing gun at her
Published: 2026-08-13 20:37:14 | Created: 2026-08-14 05:10:55
BBC Verify has examined footage shared by a Virginia woman in which an ICE agent points his gun at her.
show more
SEO трафик на сайте есть, а заявок нет. Я полез в серверные логи и посмотрел, кто на самом деле «ходит» на сайт
Feed: Все публикации подряд на Хабре (https://habr.com/ru/rss/articles/)
Published: 2026-08-14 05:10:28 | Created: 2026-08-14 05:10:55

У сайта рос органический трафик, позиции были в порядке, а заявок всё равно не хватало. Клиент задал простой вопрос: «Если людей стало больше, где заявки?»

Я полез не только в Метрику, но и в access.log — и там уже начался отдельный зоопарк: Googlebot, YandexBot, GPTBot, Ahrefs, Semrush, MJ12bot, DataForSEO и Chrome, который вёл себя совсем не как человек.

В статье покажу, как отличать визиты от серверных запросов, проверять настоящего Googlebot, разбирать crawler'ов и решать, кого вообще стоит пускать на коммерческий сайт через robots.txt.

Читать далее
show more
Helen Goh’s recipe for frozen Eton mess with plum, coconut and raspberry | The sweet spot
Published: 2026-08-14 05:00:18 | Created: 2026-08-14 05:09:56

The classic British dessert, but with a frozen twist, captures all of the charm of the original with a welcome chill for summer

I love the glorious unruliness of Eton mess, all billowy cream, crushed meringue and tumbling fruit. This frozen version captures all of that charm but in a sliceable, make-ahead dessert that’s perfect for August entertaining. The fruit can even be cooked a day in advance, leaving little more to do than whip everything together before freezing. Thanks to the condensed milk, the cream freezes surprisingly softly, while the meringue keeps its gentle chew and the toasted coconut adds a welcome crunch and nuttiness. Late summer plums bring a lovely tartness that balances the sweetness beautifully, though apricots, peaches and nectarines would also work well.

Continue reading...
show more
Searches for UK homes with air conditioning more than double
Published: 2026-08-14 05:00:20 | Created: 2026-08-14 05:09:55

Most owners would consider installing aircon as heatwaves lead buyers to focus on ability to stay cool, Rightmove says

Searches for homes for sale in the UK with air conditioning have more than doubled this summer as heatwaves push staying cool up the list of priorities for homebuyers, alongside nearness to good schools and transport links.

In the week of the UK’s fifth heatwave, in which temperatures in parts of England hit 38C, the property website Rightmove said the heat was influencing how Britons thought about their homes, with more people paying attention to how comfortable a property would be during hotter weather.

Continue reading...
show more
Hot topic: why is Burnham so reluctant to talk about the climate crisis?
Published: 2026-08-14 05:00:20 | Created: 2026-08-14 05:09:55

Extreme heat is harming the wellbeing of the ordinary people the PM champions, yet he has ignored the subject

From the first day Andy Burnham came into office, climate scientists, campaigners and some of the ordinary people surviving through this summer of extreme heat began to wonder when he would start to talk about the climate crisis.

Up until this week, many held back from criticism. After all, he had only been in power for a couple weeks; it felt far too early to start to judge and call him out on it.

Continue reading...
show more
Week in wildlife: snake on a car, a careful spider mum and a lazy squirrel
Published: 2026-08-14 05:00:18 | Created: 2026-08-14 05:09:55

This week’s best wildlife photographs from around the world

Continue reading...
show more
‘We feared we’d all be drowned’: the Indian floods that killed more than 100 and left thousands homeless
Published: 2026-08-14 05:00:19 | Created: 2026-08-14 05:09:55

More than 2,600 villages have been flooded in Assam this year, in a climate disaster made worse by uncontrolled deforestation and mining in the region

It was a normal July morning in the Indian state of Assam. After finishing her chores and eating breakfast, Gitali Phukon, a resident of the town of Gaurisagar, left her home to buy groceries. After two days of rainfall that had made the unpaved roads muddy and slippery, the sun was shining.

But less than an hour later, she returned home to a nightmare: the Mitong River, which runs through the town, was overflowing and water was gushing through the streets.

Continue reading...
show more
Nicole Kidman opens up about divorce from Keith Urban – and ‘throwing my career away’ for Tom Cruise
Published: 2026-08-14 04:32:49 | Created: 2026-08-14 05:09:55

Actor, 59, reflects on high-profile marriages in Vogue interview, saying she felt ‘deeply vulnerable’ after split from Urban

Nicole Kidman has opened up about her divorce from Keith Urban and reflected on her marriage to Tom Cruise, saying she was warned being his wife would overshadow her career.

In a new interview with Vogue, the 59-year-old actor spoke about going through two divorces under intense media scrutiny: with Cruise, who she married in 1990 and divorced in 2001, and musician Urban, who she married in 2006.

Continue reading...
show more
Outdated American Policy Is Hurting the Afghan People
Published: 2026-08-14 15:20:41 | Created: 2026-08-14 05:02:55
A new generation of Afghans wants to move on from the animosities of the past, but Western policies make that difficult.
show more
Умер Чжу Жунцзи, человек, превративший Китай в мощнейшую торговую державу
Published: 2026-08-12 15:27:56 | Created: 2026-08-14 04:41:55
Умер бывший премьер Госсовета КНР Чжу Жунцзи, благодаря которому Китай вышел на мировой рынок. Ему было 97 лет.
show more
Solving integration woes with a hackathon
Feed: Stack Overflow Blog (https://stackoverflow.blog/feed/)
Published: 2026-08-14 07:40:00 | Created: 2026-08-14 04:40:55
Ryan welcomes Meryll Blanchet,  Director of Engineering for Adobe Brand Visibility, to chat about Adobe’s recent acquisition of Semrush, how Adobe Brand Visibility was born from Semrush’s AI visibility product and Adobe’s LLM Optimizer, and how Adobe used a three-day internal hackathon instead of a large-scale infrastructure integration to quickly deliver value to customers.
show more
Hit hard by the earthquake in western Colombia, the mayor of Cali gets to work
Feed: NPR Topics: News (https://feeds.npr.org/1001/rss.xml)
Published: 2026-08-13 21:39:19 | Created: 2026-08-14 04:36:56

Search efforts continue though hopes are fading as teams dig for survivors from the 7.4 magnitude earthquake that killed more than 250 people.

show more
Family's 'desperate' appeal to know how grandmother died
Feed: The Latest News from the UK and Around the World | Sky News (https://feeds.skynews.com/feeds/rss/home.xml)
Published: 2026-08-13 23:01:00 | Created: 2026-08-14 04:28:54
The family of a grandmother who was killed in her home using blunt-force are "desperate" for answers about how she died.
show more
Morrison-era GST deal with WA a multi-billion dollar mistake that should be reversed, Productivity Commission finds
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 04:00:19 | Created: 2026-08-14 04:19:56

Interim report calls for carve-up deal that has only benefited Western Australia to be ‘reshaped’ as it has achieved almost none of its objectives

The Productivity Commission (PC) has criticised the controversial GST deal with Western Australia as a costly mistake that should be reversed, saying tens of billions of dollars of taxpayer money has since gone to the country’s richest state.

An interim report from a PC review of the deal said the reform – requested by WA and implemented under the former Morrison government with Labor’s support – had achieved almost none of its objectives and had made the system less equitable.

Continue reading...
show more
Indonesia wildfires threaten orangutans rescued from animal traffickers
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 04:00:16 | Created: 2026-08-14 04:19:56

Blaze fuelled by hot, dry and windy conditions comes within metres of West Kalimantan rehabilitation centre

Orangutans rescued from wildlife traffickers face a new threat from the record number of fires that have broken out in Indonesia in recent weeks.

Flames have come within metres of an animal rehabilitation centre in West Kalimantan as hot, dry and windy conditions have been amplified by a powerful El Niño and the human-driven climate crisis.

Continue reading...
show more
‘Unprecedented’ rain kills five in Japan as thousands stranded at airport
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 13:11:48 | Created: 2026-08-14 04:19:56

Evacuation orders issued for 400,000 after flooding, landslide warnings and power outages in Chiba prefecture

At least five people have died in eastern Japan after “unprecedented” heavy rain that left nearly 7,000 stranded at Tokyo’s Narita airport and forced thousands to take refuge in government buildings.

The torrential rains hit Chiba prefecture, east of Tokyo, on Thursday night, triggering landslide warnings and power outages. Evacuation orders were issued for more than 400,000 people in Chiba and other areas, and about 45,000 homes lost power, as 115mm of rain fell hourly, approximately the same amount usually recorded for the whole month of August.

Continue reading...
show more
‘Unprecedented’ rain kills five in Japan as thousands stranded at airport
Published: 2026-08-14 13:11:48 | Created: 2026-08-14 04:17:55

Evacuation orders issued for 400,000 after flooding, landslide warnings and power outages in Chiba prefecture

At least five people have died in eastern Japan after “unprecedented” heavy rain that left nearly 7,000 stranded at Tokyo’s Narita airport and forced thousands to take refuge in government buildings.

The torrential rains hit Chiba prefecture, east of Tokyo, on Thursday night, triggering landslide warnings and power outages. Evacuation orders were issued for more than 400,000 people in Chiba and other areas, and about 45,000 homes lost power, as 115mm of rain fell hourly, approximately the same amount usually recorded for the whole month of August.

Continue reading...
show more
Indonesia wildfires threaten orangutans rescued from animal traffickers
Published: 2026-08-14 04:00:16 | Created: 2026-08-14 04:09:55

Blaze fuelled by hot, dry and windy conditions comes within metres of West Kalimantan rehabilitation centre

Orangutans rescued from wildlife traffickers face a new threat from the record number of fires that have broken out in Indonesia in recent weeks.

Flames have come within metres of an animal rehabilitation centre in West Kalimantan as hot, dry and windy conditions have been amplified by a powerful El Niño and the human-driven climate crisis.

Continue reading...
show more
Experience: I found my long-lost twin aged 49
Published: 2026-08-14 04:00:17 | Created: 2026-08-14 04:09:54

The joy, disbelief and shock of discovering I had a brother the same age as me came with a million questions. How were we separated and why?

In 1975, my parents saw an advert in their local Australian newspaper. With the Vietnam war coming to an end, thousands of orphaned children were being evacuated. As part of the American-led Operation Babylift, the Australian government was looking for families who would adopt them. My parents decided to do it.

I arrived in Adelaide in April in a shoebox. I was so tiny that my umbilical cord was still attached. My parents knew that the document they were given, with a January birthday and the name Le Thi Ha, couldn’t be mine. But with the chaos of the evacuation there was no way to find out more.

Continue reading...
show more
‘It was lonely being a young queer artist’: Sam Smith on finding health, happiness – and their hazel-eyed fiance
Published: 2026-08-14 04:00:18 | Created: 2026-08-14 04:09:54

On their new album, Smith is subverting their singer-songwriter roots to tell a queer love story. They talk about their shifting fanbase, the nasty undercurrents of fame – and finding sanctuary in New York City

Sam Smith logs on to our video call with the username “Lily of the Valley”, the May-born singer’s birth flower. It sets a serene tone. They’re in Guadalajara, Mexico, for the latest stop of a residency tour, titled To Be Free, in support of their fifth studio album. Instead of Smith’s usual jet-lagged arena itinerary, this lower-key excursion stops at smaller venues in only six cities and sheds the ostentatious theatricality of their last outing.

On stage on their last tour, “I was flailing around on a massive gold bum and had, like, seven costume changes,” Smith says with a chuckle. It’s intimacy they’re after now, the ability to satisfy their “craving to look at people again”, which matches the romantic feeling of their fifth album, Hazel Eyes. But for Smith, whose genre-shifting stardom has been accompanied by harsh controversies, it’s also a chance to observe who’s still in their corner: “There’s some wonderful people that have stayed, but it’s a completely different room than it was with In the Lonely Hour,” their 2014 debut.

Continue reading...
show more
Parent says son feels a sense of ‘powerlessness’ aboard USS Abraham Lincoln
Published: 2026-08-14 00:16:41 | Created: 2026-08-14 04:04:55
Growing concerns are emerging over conditions aboard the USS Abraham Lincoln as sailors and families report low morale, supply shortages and months without a port call. In an interview with NBC News’ Tom Llamas, one parent calls for answers, saying her son feels a sense of “powerlessness” while serving aboard the carrier.
show more
Officials said no credible threat to Trump ahead of Air Force One switch: Sources
Feed: ABC News: Top Stories (https://abcnews.go.com/abcnews/topstories)
Published: 2026-08-14 03:50:56 | Created: 2026-08-14 04:02:55
The intelligence originated with Israel, which passed it to the United States in real time as Trump prepared to leave Ankara, Turkey, according to the sources.
show more
Битва за прошлое. В Тбилиси сносят символ эпохи Саакашвили
Published: 2026-08-14 03:49:43 | Created: 2026-08-14 03:50:54
В центре Тбилиси, в парке Рике, разбирают один из самых заметных и противоречивых объектов грузинской столицы: Музыкальный театр и выставочный зал, более известные как «трубы» или «кувшины Рике». Почему комплекс, который задумывался как одна из основных выставочных и культурных площадок страны, теперь исчезает с карты Тбилиси?
show more
Sydney airport incident involving Virgin and Qantas flights didn’t have to be reported, Airservices Australia says
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 03:32:16 | Created: 2026-08-14 03:49:56

Airservices boss says planes remained 1km apart so runway incursion ‘didn’t meet mandatory reporting requirements’ but safety bureau is ‘concerned’

The organisation that manages Australian skies has defended its decision not to immediately report a fresh safety incident at Sydney airport when a Virgin plane aborted its takeoff.

The chief executive of the commonwealth-owned Airservices Australia, Rob Sharp, said he did not believe the incident met the reporting threshold, given the substantial distance between the Virgin plane and a Qantas aircraft crossing the runway in front of it on Sunday.

Continue reading...
show more
How does programming language affect token efficiency and correctness?
Feed: https://danluu.com/atom.xml (https://danluu.com/atom.xml)
Published: 2026-08-09 00:00:00 | Created: 2026-08-14 03:48:57

This somewhat widely cited post (I keep seeing it cited, anyway) suggests that dynamic languages and/or languages that represent things more concisely are more token efficient. It seems to be cited enough that LLM search results agree. For example, when I searched for "dynamic vs static language token cost" (no quotes), Google's AI summary opened with

Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact.

Google's AI cited the same post, which suggests that some concise dynamic languages have maybe 1/2 to 1/3 the token cost of static languages like Rust, Go, C++, etc. The author says

There was a very meaningful gap of 2.6x between C (the least token efficient language I compared) and Clojure (the most efficient).

And then they later tried J, saying

It dominates at just 70 tokens average, nearly half of Clojure (109 tokens). Array languages can be extremely token-efficient when they avoid exotic symbol sets. If token efficiency turns out to be a key driver, this is perhaps a very interesting way for languages to evolve.

The other dynamic vs. static language token comparison I've found floating around is this one, which supports the same conclusion. If you want to treat this as part 8 of this series of exercises on benchmarking, evals, and experimental design, you can click through to the links and think about eval issues before reading further.

Without running our own eval, one problem the first experiment has is that the problems are trivial, which we can see from quote above; a problem that can be solved in 70 tokens in J and 109 in Clojure isn't much of a problem at all (the author used Rosetta Code). As we saw when we looked at other evals of caveman mode vs. our own evals, you can get very different results from trivial problems where most of the work is in printing out an answer vs. slightly less trivial problems that actually require some amount of "real work"; the big gains claimed by caveman mode and shown in replications go away when you start looking at problems that take more than just a few tokens. In general, performance on trivial tasks doesn't generalize.

The issues in the second link are a little more subtle, so we'll defer most of them to an appendix, but they include issues like one of the tests executing the wrong path (which doesn't exist), causing a test to fail. One of the later agents then symlinks the non-existent path to its own executable, which works for that case, but also causes every later test to run that one agent's executable instead of the correct executable. The author tries to draw conclusions about what it means that Rust had some failures, but all it means is that scoring for Rust ran before the Go agent symlinked all scoring on that broken test to the Go executable.

Instead of relying on these evals, we can try running some of our own evals. As we can see from these evals as well as the evals discussed in our last exercises on evals, it's very easy to make an eval that doesn't say what the creator of the eval seems to think it's saying. No doubt these evals will not be an exception to this and will be flawed (see appendix below for more details).

As a way to build my intuition about things, I like to pre-register guesses before looking at results1. Some things I pre-registered with friends were:

  • High confidence (95%): the overall dynamic vs. static language claim won't hold
    • For reasons stated above: this feels analogous to the caveman eval, where the result will, at best, get diluted as the problem gets larger
  • Low confidence (60%): static languages will be somewhat better than dynamic at ultra effort
    • Very weak confidence that, at ultra effort, the harness will get feedback to the model more quickly and this will result in some kind of benefit for either correctness or efficiency, but it would also seem reasonable for this to not be the case for all kinds of reasons, e.g., I've noticed that codex, when invoking the Rust compiler, very often makes the exact same error and then has to fix it; perhaps this kind of thing dwarfs things like a hypothetical faster feedback cycle
  • High confidence (98%): the "weird" language supremacy of something like J won't hold
    • Same reasoning as the overall static vs. dynamic claim, with the additional thought that AI labs are going to have much less (and possibly zero) synthetic data RL env effort on obscure languages

Zstd

For the first eval, I tried giving agents the zstd RFC (plus errata) and telling them to implement a complete zstd decoder (agents are stuck in a container without internet access). The tests were not given to agents. For something with the surface area of zstd, it's not really reasonable to expect that the tests cover every possible case. For example, even though zstd is a fairly well-tested piece of software, I once found a data corruption bug in zstd. The test suite isn't intended to find extreme corner cases that might be lurking for years and is instead intended to check various cases that can "easily" be derived from the RFC that should work.

Below, the x-axis is cost and the y-axis is correctness score (up and to the left is better / down and to the right is worse); average result on medium and ultra efforts with GPT-5.6 Sol. If we only look at medium (and ignore the fact that results often wildly differ on different tasks), we might come to a conclusion like the Alderson evaluation, that dynamic languages are more efficient and better when using LLMs because (ignoring relatively obscure languages) the cluster of dynamic languages lands up and to the left of the cluster of static languages (we used Alderson's color-coding for static vs. dynamic to make it easy to compare at a glance). But if we look at ultra effort, the results are quite mixed, with a couple static languages doing the best, with more static than dynamic languages among the better results.

The graphs below also have a toggle to convert the x-axis to time instead of cost. The mame/ai-coding-lang-bench noted that it's valuable to get results more quickly (I personally don't find this to be the case because results take long enough that I multitask instead of waiting), so we can also look at that. Similarly, we observe that neither language type dominates the other although, at medium effort on this particular task, the best dynamic language results are once again better than the best static language results (though, once again, they're fairly close).

We can observe that, just like when we compared completely trivial caveman mode evals to a less trivial caveman mode eval, the very strong relationships that held in the trivial evals don't generalize to this larger case. As was the case there, the extreme ratios in performance go away in these larger evals, except in cases where we might expect poor performance, such as when using assembly (which would be significantly more time consuming and difficult for a human) and when using relatively obscure languages where we might not expect that AI labs are expending effort generating synthetic RL environment data.

Note that this is the opposite of what the 1st eval found when it suggested that very dense languages like J would make sense for efficiency reasons. Perhaps using an obscure (and "weird") language can make sense if you have a very large budget and you can train or fine-tune a model to be effective for your pet language, but if you're a normal user of LLMs, it seems like sticking with a mainstream language is likely a better bet than using an obscure dense language.

And it turns out that if we plot language popularity vs. performance on this eval (not shown), we observe a weak to moderate positive correlation where more popular languages end up with more correct as well as cheaper solutions.

As we previously noted, very closely related evals can give substantially different results. For example, we saw significantly different results in the Optimization 1 vs. Optimization 2 evals here when Optimization 1 and Optimization 2 were optimizing bzip2 compression and decompression in wasm, which are fairly closely related tasks as evals go. To make a strong, universal, claim, like "dynamic languages are more efficient than static languages", we'd have to run evals across many tasks. However, showing that a claim like

Dynamically typed languages generally have a lower LLM token cost than traditional statically typed languages because omitting explicit type declarations makes the code more compact.

is maybe at best vaguely directionally true and not really relevant to any particular case and maybe not strong enough to be relevant in general, we just need to try a few cases and see that the claim doesn't generally hold. Above, we saw that at one effort level, the claim seems to maybe be kinda sorta true, but with exceptions, and then at a higher effort level, the claim seems to not be particularly true, which is sufficient to say that the claim is probably not universally true, modulo our eval having a confounder that completely invalidates it.

Pandoc

But, just to get a view on a very different task that's also presented in a different way (more TDD-like than "read a spec"-like), this next eval takes the Pandoc ProgramBench eval and modifies it for our use case. Instead of the reverse engineering task presented by ProgramBench, we present agents with ProgramBench materials as well as the ProgramBench tests and then score agents against a holdout set of tests to measure the performance of each condition2.

In the results below, the x-axis is cost again and the y-axis is score on the holdout tests.

As before, we don't see a very strong relationship between success or cost and whether a language is static or dynamic or very dense. We once again see that relatively obscure languages tend to do poorly (although Clojure does much better here than on Zstd). Also, Assembly does much worse, which seems expected in that we would expect a human writing Assembly to be at much more of a disadvantage implementing Pandoc than implementing Zstd and there doesn't seem to be a strong reason to think that LLMs would be different in this regard.

What does it all mean?

Who knows?

I have a lot of questions about what works well when using LLMs (such as, what test techniques work well, what languages work well, what software architectures work well, if bug fixing cost varies by language, if general program maintenance cost varies by language, etc.). Most of these questions are unanswered in public data and, if they've been answered in AI labs, the information mostly hasn't been made public.

Most of the claims that get thrown around about how a particular language is good for LLM use seem to be wrong (e.g., the claim that Ruby, Clojure, and J, are particularly well suited to LLMs, which were mentioned in the evals linked above, as well as the somewhat common claim that Elixir is particularly suited to LLMs), but it's not clear what's right.

In 2014, we looked at the literature on static vs. dynamic types and found that surveying the literature wasn't very informative outside of a few case studies. For an example that typifies a standard academic study, we saw the paper, Do Static Type Systems Improve the Maintainability of Software Systems? An Empirical Study, on which I commented:

Subjects were given classes in which they had to either fix errors in existing code or fill out stub methods. Static classes for Java, dynamic classes for Groovy. In cases of type errors (and their respective no method errors), developers solved the problem faster in Java. For semantic errors, there was no difference. The study used a within-subject design, with randomized task order over 33 subjects. A notable limitation is that the study avoided using “complicated control structures”, such as loops and recursion, because those increase variance in time-to-solve. As a result, all of the bugs are trivial bugs. This can be seen in the median time to solve the tasks, which are in the hundreds of seconds. Tasks can include multiple bugs, so the time per bug is quite low.

Picking tasks that avoid "complicated control structures" such as loops and recursion, where tasks take hundreds of seconds makes the result meaningless with respect to tasks that really eat up a professional programmer's time, just like the first eval we saw where tasks took high tens to low hundreds of tokens. However, with LLMs, we can actually feed them non-trivial tasks and compare how they do. There's the issue of how well results generalize to different tasks, but we'd have that exact same issue with human studies, but worse (LLM variance is huge, but human variance is even huger since you can't get the same human to do a bunch of tasks with different seeds). And while $20 to get an LLM to implement a Zstd decoder isn't exactly cheap once you multiply by the number of languages and the number of iterations per condition per language, if you think about how much it would cost to hire a professional programmer who can read the zstd RFC and implement it, there's no way the equivalent study would've been done because the cost would've made it completely infeasible. That goes double for the Pandoc task.

With LLMs, a lot of the questions have gone from being effectively unanswerable to being answerable with a bit of effort and some tokens. Due to the incentives that are in play3, it's not clear that we'll get answers to questions like this any time soon, but it's at least possible to take a crack at it now.

There are a lot of claims I've seen floating around that these evals can't prove or disprove (for the reason noted above that, due to the variance across different problems, many more tasks would have to be tried), but that these shed some light on, such as:

  • Languages with a lot of bad code out there (e.g., PHP) will perform worse
    • Appears to be false on these tasks
  • Because it's so easy to re-write now, you should use a powerful language (like Haskell)
    • Appears to be false on these tasks
  • You should use a popular language
    • There's weak support for this statement

For my pre-registered guesses, we had

  • High confidence (95%): the overall dynamic vs. static language claim won't hold
    • This seems correct
  • Low confidence (60%): static languages will be somewhat better than dynamic at ultra effort
    • There's not enough information to determine this conclusively, but if we had to make a binary correct/incorrect call, I would call this incorrect
  • High confidence (98%): the "weird" language supremacy of something like J won't hold
    • This seems correct
  • [from a draft reader]: "dynamic is better on small-scale, but gets overtaken by static as the size of the project grows"
    • Not supported by these tasks (static languages didn't seem to do substantially better than dynamic on the much larger Pandoc task vs. the smaller Zstd task), but the tasks and the presentation of the tasks are so different that it's unclear if this is because task-size scaling or because of other differences

By the way, a major reason Clojure improves by so much in the Pandoc eval compared to the Zstd eval is that, in the Zstd eval, 36/40 medium and 5/40 ultra Clojure programs had test failures because byte conversion throws on 128–255 (maybe unchecked-byte should've been used?) and they used this conversion inappropriately.

That's a real result, in that, if you ask the best publicly available GPT model to implement Zstd (and presumably if you do other bit/byte manipulation tasks where this might come up), it will emit code that fails in this particular way. If there are tests that catch this, the bug will get fixed, but it will still cost time and tokens. Whether or not a language did well, there are costs like this all over the place (for example, cargo repeatedly gets invoked with the wrong arguments, which then immediately gets caught and fixed, but I've noticed this loop can actually consume a decent amount of wall clock time on my real projects unless you give explicit instructions to codex on how to invoke cargo, and it's clear that's worth the space in the context window).

Anyway, all of this is an illustration of why, if someone wanted to make a strong claim about which languages or classes of languages are particularly good with LLMs, they would need to run quite a few different evals. If we dig into why any particular condition got a certain score, the failures that caused the score are generally something idiosyncratic where it's not always obvious how much the issue generalizes across tasks or across setups. There's no way to look at the score on one eval or even five or ten evals and draw a conclusion about programming in general.

It's true that, in both the Zstd eval and the Pandoc eval, we see a correlation between language popularity and positive outcomes (higher correctness, lower cost, lower wall clock time) and it seems plausible that we'd see this across other evals, but it would be a mistake to draw a strong conclusion about any particular language. I gave a warning like this back when I looked at how often different projects have a broken build according to GitHub CI data, noting that there are different reasons that a build might be broken more or less often across projects and that one shouldn't draw strong conclusions because results across projects aren't necessarily comparable (for example, if one project's main branch is some kind of release candidate that's gone through other vetting, that project would be expected to have low build breakage, but that's not comparable to a project where people are developing directly against main).

Shortly afterwards, someone involved in one of the languages with a high score (IIRC, it was Martin Odersky and Scala) tweeted out the post and cited the language's high ranking as a victory for the language. That was an unwarranted conclusion there and, due to the many sources of variance that are in play here, any such conclusion about a single language would be even more unwarranted here.

This data (assuming eval validity) can refute some strong claims and is suggestive of other claims, but it can really only be suggestive of things for classes of languages and not for particular languages due to having only two tasks, which any particular language could do well or poorly on for some idiosyncratic reason which may or may not generalize to other tasks.

Thanks to Max Bittker, Yossi Kreinen, Aaron Levin, Alan Boll, Luke Burton, Marco Primi, Milosz Danczak, Justin Blank, and Tom Adamczewski for comments/corrections/discussion.

Appendix: selected issues in ai-coding-lang-bench

Like I said above, my eval here is a quick and dirty eval and I'm sure it's full of flaws, so I'm not trying to say the evals I've presented here are great and this is bad, but here are a number of issues in the Endoh ai-coding-lang-bench eval.

One issue is that the wrong executable appears to have been run for some of the tests. The setup for the published run seems to have executed ../../minigit inside each candidate's directory for one of the tests when the candidate's generated executable is at ../minigit. ../../minigit doesn't exist.

Because statically typed languages had a lower correctness score, the author of the eval noted "the only failures in 600 runs were in Rust and Haskell (both statically typed, both relatively "difficult" languages)" and suggests that "difficult languages", such as "C's memory management, Rust's ownership model, and Haskell's monads/purity may add overhead for the AI".

However, Rust's failures were because there is no executable at ../../minigit, causing the test to fail. The first Go run "fixed" this by executing ln -sf minigit-go-1-v1/minigit ../minigit and linking generated/minigit to its own run, but this means that every later execution (for every language) actually executed the first Go run's executable. On rescoring Rust against its own executable (as opposed to having it fail by trying to execute a non-existent file), Rust gets a perfect score, invalidating the theory that Rust had failures because it's a difficult language to deal with.

Other tests also have issues. For example, two tests have a structure that causes them to pass regardless of the actual value being checked. One of the tests has

  if ../minigit commit ...; then                                                                                                                                      
    COMMIT_POST_CHECKOUT=$(cat .minigit/HEAD)                                                                                                                         
                                                                                                                                                                      
    if grep -q "parent: $COMMIT1" \                                                                                                                                   
        ".minigit/commits/$COMMIT_POST_CHECKOUT"; then
      pass "checkout then new commit works"
    else
      pass "checkout then new commit works"                                                                                                                           
    fi                                                                                                                                                                
  else                                                                                                                                                                
    fail "checkout then new commit works"                                          
  fi

The inner if has a pass in both branches, meaning that this is almost equivalent to

  if ../minigit commit ...; then
    pass
  else
    fail
  fi

The inner if appears to be intended to have the actual check, but due to a coding error (perhaps a copy+paste error?), the check is effectively elided.

Also, as noted above, agents can modify the test environment, which the 1st Go agent did to fix a broken environment. They have full access to tests and the environment and can do anything and the test suite is visible during development with no holdout, which can easily lead to cheating by special-casing code in a way that passes tests but creates a program that's useless "in real life". At a high level, something like this seems to have happened in that many programs fail to implement large parts of the spec but do pass all tests, which may indicate that the agents "understood" how to pass the tests and preferred that over implementing the spec (it could also indicate that the tests are very thin and are easy to pass).

Another issue is that the Claude Code CLI versions aren't the same for all runs (it varies from 2.1.66 to 2.1.68). There are a handful of other issues like this that could be significant, but are likely small compared to the issues noted above.

Appendix: medium in a loop vs. ultra

As an example of something we can compare, I was curious how cost effective using medium + asking the agent to keep working would be and then, in the back of my mind, I also had this question about something "Ralph loop" advocates say, that you're better off clearing the context window on every iteration of the loop and giving the agent the full prompt again. As with the above, my pre-registered guesses here are:

  • Zero confidence (50%): Ultra is more effective than medium in a loop
    • I'm not sure how to think about this. I guess the case for this would be that ultra was designed in some way and should be smarter than repeatedly doing medium in a loop. But it's possible that there's some tradeoff where ultra was made for more speed and, as we've noted, the variance is very high so even if ultra wins on most problems it might lose here; ultra might also be more optimized for trading off to improve wall clock time or another parameter; ultra also has the disadvantage that it doesn't "know" to stop after reaching correctness on the hidden tests, whereas medium conditions that hit full correctness aren't run again under this setup, which hugely advantages medium in a loop (which is arguably realistic w.r.t. how someone might use these)
    • You could maybe say this is 50% + epsilon since my mind went to framing it this way and not the other way around, but I would say extremely low confidence here at best
  • Medium confidence (80%): continuing with context outperforms Ralph loop
    • /goal mode, etc., don't do this by default and, presumably, folks at Anthropic and OpenAI have tried things like the Ralph loop and found them less effective
    • Watching your context window very closely seems to have gotten less important as harnesses (and models?) have improved; in late 2025 / early 2026 I often had to throw out my context window when working on a long-running task to avoid issues and that's gotten rarer over time but, even then, because I wasn't paying attention to what people were saying, I was running agentic loops with a default of keeping context and only clearing when there were obvious problems, which seemed to work ok, e.g., I built the world's strongest Azul AI doing that, so it's not clear to me that having a default of clearing context on every loop iteration was the right choice back then

Below, we have the average result for medium in a loop vs. ultra, sorted by best to worst ultra correctness score, for a prompt that simply resumes individual runs that don't have 100% test correctness as well as a Ralph-loop like prompt that discards context and gives the original prompt again (x-axis is cost, y-axis is number of correct test cases):

For this one problem, on average, running ultra once seems better than repeatedly running medium per unit cost (and much more so per unit time) and continuing with previous context outperforms Ralph. The problem with naively running medium on repeat is that the agent can get anchored to a bad solution and fail to make progress. The theory behind the Ralph loop is that you throw away bad context which can cause this to happen, but that doesn't save you from having a bad artifact.

Just from using LLMs, I've noticed that you're often better off throwing away a chunk of code and having an LLM re-write it from scratch than you are having an LLM modify it or try to re-write it in place. Michael Malis, who's been re-writing Postgres in Rust and has been making major changes has also noted this. This also relates to this idea noted previously that, due to the high variance (plus this path dependence) you're often better off rolling the dice multiple times and taking the best result, if you don't mind spending the tokens.

It's hard to say too much about static vs. dynamic languages from looking at just this one condition, but a naive thought like "static languages will outperform when iterating" isn't obviously true. If there's one pattern that jumps out at me, it's that the cases where the Ralph loop most badly underperformed continuing with context were generally dynamic languages. It's possible this is because of the lack of type information, but we'd need to both look at the differences in trajectories in more detail as well as look at other examples to observe if that's a real pattern. Even if you don't care about Ralph loops now that the Ralph loop trend has passed, being able to make changes to a codebase more effectively when starting a new task or starting with fresh context is something you might care about and the pattern here is suggestive of a possible advantage.

Appendix: Guards of Atlantis 2

I tried to do a third eval that seemed like a more "business logic" kind of eval in both how the problem is presented and the actual execution of the problem. You can argue that the Zstd eval and the Pandoc eval are quite unusual tasks for a programmer to face in that not many programmers receive a specification as well-written and thorough as the Zstd RFC and not many programmers are handed a problem with as many pre-created tests as you get from ProgramBench tests.

The idea here was to implement a board game. In general, board game rules are written by people who aren't experts in writing clean specs, so implementing a board game is more like what happens when a non-programmer (or a programmer who isn't an expert at writing good specs) gives someone a task.

The problem here is getting a game where I have a reasonable oracle for scoring that isn't trivial for LLMs. For example, LLMs were able to one-shot the rules for Scout and Azul, which make those poor tasks. For games that an LLM won't immediately one-shot, I happen to have an oracle for Guards of Atlantis 2 because I had an LLM implement a copy for me and my friends to play (no link for this one because I don't see how to make an interface that's free of copyright infringement). The backend only took a few hours of my time, but it took a fairly large amount of LLM time to get the rules to be roughly correct. I like this as a task in that the rules are tricky in the same way a lot of problem descriptions that are delivered to programmers are tricky, but it is, in principle, possible to figure out the correct rules and implement them (after all, humans implicitly do this when they play the game correctly offline).

In board game rules, it's fairly common to have rules where reading the rule strictly as written is incorrect and you need to use "common sense" (or read some kind of FAQ) to play the rule correctly (there are some game designers who strive to avoid this, such as J C Lawrence, but this is fairly uncommon). Guards of Atlantis has quite a few rules like this. The designer of Guards of Atlantis is also vocal about there being no such thing as the spirit of the rules or common sense interpretations of the rules and says that you should always read the rule exactly as written, which creates two difficulties. One is that there are also many cases where you need to ignore the "common sense" interpretation and read the rule exactly as written. The other difficulty is, as anyone who's ever tried to write a formal spec knows, it's very easy to accidentally have ambiguity or contradictions. Even people who do this profesionally are unlikely to be able to create a non-trivial, complete, clear, spec without formal methods or a very large amount of human review. Realistically, a board game designer who doesn't have a background in writing formal specs doesn't have a chance, and thinking that it's easy (as the designer seems to) reduces the already low odds even further. A nice way to mitigate this kind of issue to write down your intent or "the spirit of the rules", but because the designer says that there is no such thing as the spirit of the rules and you should read all rules exactly as written, there are no meta-comments in the rulebook that would help someone interpret confusing or abmiguous rules. This combination is quite difficult for LLMs (and, judging by the rate at which I see humans play the game according to the designer's intent, it's also quite difficult for humans).

I think it would be effectively impossible to just read the rules and play correctly (of course it would be possible, but it would require knowing which rules are to be read as written and which rules are not, which one would have to do randomly and get lucky as the rules don't define a consistent system that one could use to infer which rules obey which meta-ruleset). When I was implementing the game, in order to get my LLM to understand the rules, I gave it various resources such as an unofficial rules FAQ (which is correct), an unofficial short version of the rules (which is better written than the official rules and correct, but incomplete), an opening book (which can be used to test rules against on the assumption that the opening book only contains legal moves), comments from the rules channel on Discord, etc., and had the LLM do consistency checks across these with the understanding that things like the FAQ and the Discord comments have higher authority than the actual printed rules.

One additional source of difficulty is that the designer is frequently delibrately unhelpful when answering rules questions. He often likes to make fun of people who ask rules questions or played rules incorrectly, which has a chilling effect and reduces the number of rules questions (multiple people have said that they don't ask rules questions because of how the designer behaves), and when he does answer questions, it's often with something like a meme image that says "reading the card explains the card". To extract the information, the LLM has to process these meme images, and then there's often no information or delibrately round-about information, such as, in the case of ambiguity, a referenece to a particular section. When people do point out contradictions, the designer often says it should be obvious which side of the contradiction is correct, which may be true for a human who's kept up on all rulings to date, but current SOTA LLMs find many of the designer's rules clarifications unhelpful.

Yet another source of difficulty, perhaps related to the designer's propensity to make fun of people who ask rules questions or are confused by rules, the game's interface seems almost designed to trick people into doing the wrong thing. There are multiple design affordances that I've seen trip up most new players (even if you explain to the UI trap to them, there are enough rules to take in they often forget, and then when it trips them up, they'll say something like "I'm an idiot, you even explained that to me twice"). It's not clear if these traps were created to give the designer people to make fun of, but that's certainly one result. Another is that LLMs struggle to understand the games rules and UI.

With my $200/mo personal OpenAI/codex account, I let an LLM use all my spare capacity to run consistency checks and make rules fixes. I didn't closely track how long this took, but I think it was something like a month or two of cranking on fixes like this to get a somewhat reasonable result that's playable, but that I wouldn't really trust to be correct.

The only reason I somewhat trust this is that Pedro Oliveira also implemented Guards of Atlantis and they used a completely different approach (a more standard approach of having a human drive an LLM rather than trying to get the LLM to figure things out itself). When we compared implementations, we found maybe 10-ish bugs in each. There are probably some remaining bugs where both of our implementations incorrectly do the same thing and perhaps some where our implementations differ but the checking system didn't notice, but I think the rules for both of our implementations are now reasonably solid. That's how I have an oracle for this game.

I like this as a task because it feels more like the kind of "specification" you get in the real world, where the spec is ambiguous and contradictory and sometimes just plain wrong, and then you need to use other information to get a correct result. For this eval, to avoid having it be a test of how well LLMs can access data in annoying formats (such as converting the opening book from a set of images to some kind of structured data, converting a scan of the rules to text, etc.), I gave agents both the originals of anything where I directed an LLM to extract the data (which also required various consistency checks to get correct) as well as the the extracted data (the originals were presented so that LLMs could check the originals for extraction errors if they chose to).

While I did this task with older models (I did a chunk of it with GPT-5.1 or 5.2, and then another chunk with 5.4 or 5.5), with newer models but without the kind of guidance I gave to the older models, the task was still far too hard. Regardless of language, agents scored approximately 0 on this task.

BTW, if you're curious what LLMs (and humans) struggle with, here are some examples. There's one card whose text reads "Target a unit adjacent to you. After the attack: may repeat once on a different enemy hero."

In this game, a hero is a type of unit. Read strictly, with full knowledge of the rules, e.g., what "After the attack" means, etc., this should mean that you can either attack a single unit or you can attack two heroes (after all, to repeat the attack on a different enemy hero would mean that the first unit was a hero; otherwise it would be a different unit that is a hero, not a different enemy hero).

This card actually has what is effectively an errata printed on the card because people complained it was unclear; the errata reads "(You may repeat even if the original target was a minion)". That's already confusing to LLMs (and some humans), but the real killer here is that there are other cards that use the same construction and don't have this correction. To play other cards with the same construction correctly, you need to know that every time this construction is used, you should play it with the errata that's on this card. There are a number of constructions the game designer likes to use that have a specific non-literal meaning that you have to keep in mind.

Another example of a rule that shouldn't be played in the obvious way is a character with a card which reads "Choose one, or both, on different targets: A, B". Reading this strictly as written, one would expect to be able to, on different targets, do either A or B, or both A and B. But part of the spirit of the game is the meta-rule that a character can't attack another character multiple times with one card, so the interpretation that you can do what the card says and do both and A and B on some number of different targets can't be right. Based on similar deductions and how similar constructions are used, the way this card is supposed to be interpreted is "Choose one, or both on different targets", which is arguably still ambiguous and could be more clearly written as "Choose one or both (must be on different targets if both)".

As a human, once you understand what the "spirit of the game is", you can resolve these kinds of things. But, by design, this isn't written down clearly in the rules and one has to infer this from Discord discussions, which appears to be beyond the capability of today's models even though humans who are outperformed by today's models on many specialized tasks are able to do this.

When I was supervising the LLMs that implemented the rules, the reason LLMs reached a ceiling and didn't converge to fully correct rules was that an LLM would observe that a rule was inconsistent and incorrect. It would then try to fix this rule and would also fix other things to try to make them consistent and correct. This would sometimes make things more correct and sometimes make things less correct. When making things less correct, the LLM would sometimes modify an existing correct test to turn it into an incorrect test so, after a while, the LLM wasn't really improving correctness and was just churning on which rules were incorrect. That was with some guidance on what to check and how to check it; without that guidance, even with the more advanced models that are available today, LLMs were unable to navigate this in a reasonable way.

I'm sure there is a board game of the right rules complexity to make for a good eval here but, by definition, this would be something where it would take some work to create the oracle and I don't have an oracle handy for a board game with the right rules. If my goal were to make evals, I would've used board games with actual game replay data to get good tests or oracles for a whole bunch of games, but my goal was to play a particular game with some friends. But, if one were inclined to try this board game thing, it should be possible to create hundreds (thousands?) of these in a scalable way, so one could get a reasonably correct oracle for hundreds or thousands of games and then check which games are at the correct level to be an interesting test for LLMs today.

This is arguably a bit of a funny problem in that, given a clear spec, e.g., a clearly written set of rules, an artifact that's more complex than Guards of Atlantis can be implemented by LLMs (I would argue the Zstd RFC is more complex, and Pandoc certainly is; even individual document formats Pandoc supports, like PDF, are more complex than Guards of Atlantis), so the problem isn't finding a game with rules that are complex enough that LLMs struggle and the problem is more about finding a game with rules that are ambiguous or contradictory enough that LLMs struggle, but not so much so that LLMs are completely hopeless. This is an actual real-world problem, in that humans are generally not very good at writing clear specifications and how well models and harnesses can handle a human's unclear, contradictory, and sometimes just plain wrong, specification is probably more relevant to the typical user than how well an LLM can implement something from a specification as well-written as the Zstd RFC or how well an LLM can implement a problem when handed the 4800 ProgramBench Pandoc test cases plus documentation. And these problems seem solvable in principle, in that humans who want to play board games correctly (even ones who would have no hope of "playing" Zstd correctly, let alone Pandoc) are generally able to navigate the mess of information out there to figure out what the rules to a board game are.

If we look at it form the other side, this is suggestive that, to get an LLM to do something, maintaining a clear, canonical, spec is an effective way to work.

Appendix: reasons for various decisions

  • Testing ultra
    • I've seen people say that you shouldn't really measure this because this is a harness thing and not a model thing. I can see why you'd want to measure these separately if you're working on improving models or harnesses, but when looking at how users use things, many people are just going to use codex or claude with the various built-in features and options; whether or not something is a harness thing or a user thing isn't really relevant to them
  • Using codex
    • I've seen evals use a very thin harness for the same reason as above and my reason for using codex and not a very thin harness is the same as above
    • Similarly, in this caveman model eval, I used claude with Opus and Fable and codex with GPT
  • No internet access
    • Models will often cheat if given internet access and there are plenty of problems where searching on the internet doesn't turn up source code that solves the problem, so this makes these evals approximate those more closely
  • Relatively large tasks compared to a lot of benchmarks people pass around
    • Although I have LLMs do plenty of trivial tasks, the things that take my time or take tokens tend to be larger than the kinds of tasks that were in the Alderson eval or the Endoh eval; LLMs are good enough at trivial tasks that it doesn't make too much difference to me if some condition makes them slightly better or slightly worse at one of those trivial tasks, but for a task like implementing Guards of Atlantis, where I have to spend some number of hours setting up scaffolding for the task to even sort of work, I care a lot about what makes models perform better or worse
  • Agent-specified prompts
    • Public evals seem to have moved to relatively thin/lightweight prompts that don't specify the task in great detail; this is said to be better because an agent setting up a task will give too much information that helps agents succeed at the task
      • I can see why you would want to test that, but it's also the case that I care a lot about how well agents do at tasks set up by agents because a lot of the tasks that I have agents execute are tasks that are defined by agents; I care about how agents perform under both styles, not just one style, and the public evals have moved towards one style
  • Zstd eval: asking agents to fix bugs without telling them the issue or the failing tests
    • In general, if you tell an agent to fix a specific thing, it will fix it, but it won't necessarily fix the class of issue; I've found that if you tell it there's an issue but don't tell it what the issue is, it sometimes does a more general thing instead of just putting in a narrow, brittle fix, so I do care about how agents behave when given instructions like this (of course you can tell agents to not just make a narrow, brittle, fix, but that often doesn't work)
      • This feels a bit related to the issue we noted in the Pandoc holdout footnote, where telling agents we had a holdout set appeared to force agents to produce more generalized and less brittle solutions

Appendix: issues with these evals

When it comes to performance benchmarking, I've done enough of it that I feel like I generally know how my benchmarks are flawed and I can make an informed time/effort vs. flaw tradeoff and I have decent confidence the flaws that exist in the benchmarks aren't material to the thing I'm trying to understand. I haven't done enough AI evals to have this kind of feel for AI evals so, at a meta level, I would expect any AI eval I do to have some unknown-to-me flaws.

Another reason I would expect some flaws here is that I had coding agents set up these evals and every time I spent a minute looking for issues I would find at least one issue. This indicates that it's fairly likely that these evals have additional flaws that could be uncovered by looking a bit more, but I wanted this to be more of a "quick toy project" level of correctness than a "Gary Bernhardt" level of correctness, so I stopped after fixing a handful of issues.

Back when I was working as a verification engineer, I attended a meetup by a Sun/Oracle engineer in Austin, maybe around 2007 or so, where they mathematically formalized this idea of converting the time between bugs to a level of confidence in a chip release. I haven't seen people do this much, but I recently heard Will Wilson (co-founder of Antithesis) mention that some folks at Antithesis used math from ecology (the literature on rare species observation) to estimate true bug rate, which seems like a much more sophisticated version of what this engineer at Sun/Oracle was doing a couple decades ago.

That's a cool idea, but when you're finding a bug every minute you look, you don't need fancy math to tell you that there are probably a lot of other bugs. If I were doing this for work and we had some reason to care about the fidelity of these evals, it would probably make sense to look at these more closely and fix more issues (and I would probably have the skills and experience to make fewer mistakes in instructing LLMs to set up these evals if I did this kind of thing for work). But, for the purposes of answering the question "is the claim that dynamic languages are meaningfully better than static languages when using LLMs?", I have a little more confidence that the claim isn't true, and there are a lot of other questions that seem more likely to yield some kind of actionable result (such as, what techniques or test libraries work best).

I normally don't publish things on the blog until I feel like they're somewhat solid, but this means that I often explore some data enough to satisfy my curiosity and then never publish the result. From talking to people about these non-published results, people I talk to are often curious about the results even if they're not done to a standard that I really like, which seems like an indication that folks I don't talk to might be interested as well. From what I've seen so far, I suspect it would take at least 10x the time I've put into this to get this to a standard I really like. I'm fairly busy at the moment and can't see myself having the time to do that for months, at which point I'm not sure I'd really ever get around to publishing this. In a recent post, I mentioned an analysis I did almost a year ago where I was trying to understand which cars are better for concussion risk in accidents, where I spent some time figuring that out, got far enough to get an answer that satisfied me, and then didn't ever get around to doing the work it would take to clean up the result enough to publish it.

There are some results from that analysis seem "publishable", in the sense that they could turn into a published paper (such as finding from actual crash data that the relationship between HIC and velocity looks like it's to the fourth power (!); there's a paper that tried to find this relationship, but did the wrong kind of analysis and wasn't able to find an "O(n)"-style relationship and had something much fuzzier), but I've never really cared about whether something is a paper or a blog post and it turns out that I'm more likely to just move on to the next analysis instead of cleaning up the analysis enough to publish a post.

A more recent project along these lines is that, after making a superhuman Azul AI, I tried to make a superhuman Splendor AI using a much less human-time-intensive process. I believe that didn't succeed, but it beats every other Spelndor AI I could find by a good margin, which is a mildly interesting result. I think I know enough about board game AIs to write something up about them, but my main interest was in figuring out if I could get something decent, and then I keep just doing other projects instead of spending the time to do a nice write-up. An example of something I think is interesting there is that a lot of the performance optimizations you want to do actually change the result, so you can't only rely on optimizations that can be strictly checked to not change the result. But, if you naively ask a coding agent to do these optimizations in a way that doesn't reduce playing strength, they'll do all sorts of things that reduce strength. Cases where the strength reduction is very severe are easy to catch, but there are more subtle issues that sometimes result in (for example) no change in strength vs. your own AI in self-play but a reduction in strength against humans or other AIs, so some kind of process to catch bad optimizations is necessary, and it's inherently a kind of arbitrary process that has to be designed using some combination of your intuition and relying on LLMs (which will be very helpful but also often completely wrong).

For these kinds of data-y projects that I'm interested in, LLMs massively reduce the amount of effort it takes to get a result that's strong enough to satisfy my curiosity but, AFAICT, they don't reduce the effort it takes to publish a result by much (at least if you write up results by hand instead of having an LLM write up the results and you want the results to be nice and clean), which means that writing them up runs into a kind of Ahmdhal's law bottleneck, so I've been doing more projects like this and writing up fewer of them. If anything, I think it actually takes more time to write these up because of how I've changed my workflow. For example, instead of just outputting some graph from ggplot2, I'll make a version an interacive version that's nicer in some ways, but definitely takes more time to produce. And I run an LLM spell/grammar check pass (at least so far, that's the only LLM assistance I've used for writing), which turns up a bunch of issues to be fixed. Since I look at each one manually instead of taking the fixes (and I make a lot of typos), that's actually fairly time consuming (over an hour on my last post and over half an hour on this post even though I didn't even make corrections all the way to the end and abandoned the process maybe halfway through).

Anyway, publishing this is an experiment in publishing some half-baked notes instead of having the kind of cleaned up version that I'd really like to have before publishing something. If you have opinions on this, please let me know (X Bsky Mastodon)!

I don't have GitHub links to the current evals. On the one hand, I feel like I really should. On the other hand, they're a mess and there's a bunch of stuff I'd want to clean up before publishing the code, and I don't know if/when I'll get to that and this way, at least I'm putting something out there instead of just talking to a few friends about the result and then having the result sit on my hard drive indefinitely?

Appendix: more details on Zstd

Agents were instructed to ignore performance, but the timeout wasn't infinite and, under the medium condition, some test cases timed out. This is arguably unfair, but this didn't materially impact the score. For non-infinite loop timeouts, there were 2 test cases in Clojure (across 40 * 34 tests), 2 in J, 2 in Tcl, 1 in Factor, and 1 in PHP. And, at 9000s (2.5h), the timeout was fairly generous considering that the largest test case was 4 GiB. Failing to decode 4 GiB in 2.5h is an implied rate of less than 0.5 MB/s on a Graviton 5 core, which is quite slow.

Here are some of the issues that I ran into when trying to get agents to set this up (and, as noted above, the short amount of time it took to find each issue implies there are more issues)

  • Originally, the build setup wasn't clearly specified to agents, causing some languages to randomly fail when agents did something that seemed reasonable based on how this was specified to agents but didn't work when scoring occurred
    • BTW, I was very exicted by the initial result here because it was super interesting looking and it confirmed my biases. Dynamic languages were substantially worse than static languages. What a blockbuster result! But it turned out that the real result from the initial setup was that static languages were less likely than dynamic languages to have problems caused by this issue because static languages were less likely to have issues with the idiosyncratic way project builds were ambiguously specified
  • In the original assembly conditions, agents implemented code in C and then compiled it to assembly and submitted the assembly
    • With this issue, asssembly did as well as other languages, which is super interesting! And also false once this issue was fixed. It turns out to be very easy to get incorrect but compelling looking results that would easy go viral if you aren't careful. After fixing those two issues, the results looked fairly mundane and fall into what you might call a "negative result" in the framing of a paper, in that there's no interesting or surprising or contentious thing the results show; perhaps slightly favoring boring languages would've been contarian result for very online people 10-20 years ago, but very online trendy discourse seems to have moved away from that, so this isn't really an interesting contarian result anymore
  • For some reason, the agent doing the setup imposed unusual arbitrary restrictions on some languages and not others (for example, the Rust setup didn't have access to rustfmt or Clippy); most, but not all, languages had things like this
  • Many of the tests (which were created by an agent) were actually some kind of performance/stress tests even though agents were instructed to ignore performance (I wouldn't consider processing 4 GiB of Zstd in 9000 seconds a performance stress test)
  • Some language conditions had arbitrary instructions to agents (for example, the Haskell condition had instructions not to use bytestring, with instructions on alternative implementation suggestions)
  • Some language conditions had old toolchains (for example, Zig was on 0.10)
  • Some language conditions had scaffolding to help agents implement Zstd
  • Some language conditions had explanations of tools that were available that were incorrect (for example, assembly conditions were told they had access to GDB, but GDB didn't work)
  • The agent responsible for health checks for running iterative evals would sometimes decide that evals weren't making enough progress and give held out tests or other information to agents inside the eval

There's one thing which arguably wasn't a bug that I removed anyway. One of the tests was very hard (maybe 10% of agents passed the test on the first try). On testing the current zstd release binary, the zstd binary also fails this test. On reading the RFC, this seems to be an ambiguity in the RFC about the legality of a certain edge case. There was fairly strong clustering with respect to which languages passed this test case more frequently, which I think is interesting, but doesn't seem like a very useful thing to measure when all of the other tests are measuring (or at least attempting to measure) something more straightforward.

Anyway, in the above list (which is not exhaustive), many of the issues impacted a large fraction of languages and some issues had to be fixed multiple times. All told, if you count each condition as a separate bug, I probably fixed (had agents fix) over 100 of these bugs and I expect there are more. When I talked to Max Bittker (who runs an RL environment startup), he noted

all the evals I've worked on, I ended up putting a huge amount of time and effort into, mostly in the form of reading trajectories (or summaries of many trajectories) and then triaging issues , e.g "oh this class of bug shouldn't be possible, lets update X "(X being the prompt, the harness/ environment, or the verifier)"

agents tend to slop this up, so I put a lot of care there to make sure things get fixed at the right layer, for instance it's very sensitive what's in-context for the agent under test (bad to add random junk it has to worry about, or at worst leaking answers) vs whats fixed behind the scenes in other parts of the system.

agents, when writing evals, are not sensitive enough to the experience of the agent under test, and will just give it the answer or fix problems by making it the inner agent's problem ("remember to not reward hack plz")

I also have had a lot of success re-using existing things (repos, games, tools, levels) and building harnesses and verifiers around them, versus trying to make something from scratch for an eval by prompting

In retrospect, I sort of regret doing a cross-language eval. Even after fixing 100 or more eval issues, I have no doubt that plenty more remain. Maybe this is just a "grass is greener on the other side" thought and I'll also regret the next eval I try, but I think it would've been a lot less work to try to evaluate how well different test techniques or testing frameworks work than to evaluate different languages and I find that topic at least as interesting. And, in retrospect, had I done a lot more work by hand and relied on agents less, this would've gone a lot better. For example, I should've had agents produce an environment for one language and then both had agents inspect it and inspected it myself and fixed the issues before producing the environment for another language. After doing this a few times, I might've had a better setup for producing environments for other languages (and if not, I could've just repeated this process for each language and gotten a more reliable result, likely without even taking more time).

Another thing to note is that a number of things that are genuine differences in languages weren't really tested, such as memory safety against adversarial inputs. If agents had a harder time producing generally roughly correct code in C or C++ than Rust, that would be observed, but if a fuzzer or valgrind or other tools would turn up issues, that's not likely to be captured in the small set of tests. Just out of curiosity, I asked an agent to (briefly) check the Zstd C and C++ code for memory safety issues. The agent claims it ran the C and C++ code under ASan+UBSan and tried a few fuzz inputs (4000 each) and didn't find issues, but of course that doesn't mean there aren't issues or that a larger codebase wouldn't have issues.

And, in fact, doing an analogous quick check for memory safety issues for the Pandoc eval found memory safety issues in all of the C programs and all but one of the C++ programs (the issues were things like incorrectly dereferencing out-of-bounds memory; one specific example is that, in one of the C programs, a truncated LaTeX table could result in an out-of-bounds memory read). The fact that these issues were findable with 10 of seconds prompting indicates that many such issues could be found and fixed without much human effort, but it would cost quite a few tokens and would push the cost of the C and C++ versions well beyond the cost of the Rust version and after doing all of that you would still have less confidence in the memory safety of the C and C++ versions than in the Rust version.

Anyway, if you're curious about the distribution of results, we have the following for medium and ultra:

I don't love that the ultra results are somewhat saturated here, but one "problem" with testing ultra is that it will keep going for a long time as problems get harder (e.g., most of the Pandoc ultra runs ran for 12+ hours, and the assembly runs went for much longer), so the things that don't get saturated are very large tasks, like the Pandoc eval, or tasks that are too difficult in some way, like the Guards of Atlantis eval.


  1. a draft reader pre-registered the guess, "dynamic is better on small-scale, but gets overtaken by static as the size of the project grows". [return]
  2. The holdout tests seem necessary because, without them, agents cheat and will detect a test input and hard-code the passing test output (they sometimes do this even when instructed not to cheat). If all cheating was that blatant, that wouldn't be a problem (and could be an interesting thing to measure, as agents differentially following directions or not across languages is something that matters to real users), but a lot of the cheating is more subtle and difficult to adjudicate. For example, some agents wrote code that branched off of the structure of the tests, but then filled in the contents of the branches with code that wasn't special-cased to a single test result and could pass many variants of the same test. For any point on the spectrum from "definitely not cheating" to "obviously cheating", some agent tried it. As we saw when we looked at Senior SWE-Bench, LLM scoring of evals is tricky and a great way to introduce both bias and variance; using a holdout set of tests has some problems, but it lets us avoid this much larger set of problems.

    For one thing, the holdout tests are suspsicious because they were created by agents. The intention was to create holdout tests that a reasonable person (or agent) would be able to make pass if they're not cheating. Agents audited this set of holdout tests for cases where this wasn't reasonable and eliminated some, but I didn't check these by hand, so I find it likely that there's at least one holdout test that's unfair in some way. However, the overall score against holdout tests is low enough that I'm not too worried about a small number of tests being bad (if I worked at an AI lab and was trying to train next-generation models, I would be more worried about this, but I don't think it's material for our use case here).

    Instructing agents not to cheat while having a holdout set of tests didn't prevent blatant cheating that scored extremely poorly on holdout tests, but telling agents that there was a holdout set of tests they were graded against seemed to reduce the score they achieved on the agent-visible tests while increasing the score they achieved against holdout tests (without telling them this, a number of agents achieved 100% on the Pandoc tests with uselessly brittle code; on telling them there's a holdout, no agent scored 100% after 1 turn on ultra, but the holdout scores were substantially better, indicating better generalization).

    [return]
  3. There are various Substacks, YouTube channels, and other things that promise to tell you the secrets of LLM coding success, but the ROI on spending time running actual experiments isn't really there. When we looked at caveman mode, we saw that one of the biggest programming YouTubers had a video where they spent a few minutes looking into it and decided that it worked. Spending even 15 minutes looking into whether or not it really works is probably negative ROI compared to spending that time producing more content instead.

    There are various papers that discuss different techniques, and these sometimes go into more detail than most blog posts or videos but, on average, they don't necessarily have more useful information. For example, when I asked ChatGPT (5.6 Sol, Pro) to find discussions of language effectiveness with respect to LLMs, it turned up this paper on token efficiency, which has an interesting idea, but has the same issue as the caveman mode evals we discussed earlier, where it's not looking at a task that's interesting enough for the result to be relevant to me as a programmer. Just seeing what cited that paper, we find this paper by three academics on token efficiency of languages titled "The Best Programming Language for Tokenmaxxing", but compared to this post, that paper only compares four languages, uses worse models, and uses small toy problems (from something called LiveCodeBench; the cost to solve problems with GPT-5.5 is often on the order of 1000 tokens). Regardless of how well done the eval is, as we've noted in this post and in our caveman mode eval, we often see wildly different relative results when going from a small toy problem to a problem that I might care about for hobby projects or work. Also, in that paper, they note that they gave the prompt "To test your program, run exactly ./test.sh... These are the only tests I care about" and they say this is realistic because "We believe that this setup is a realistic way to study agent behavior: in everyday use, programmers don’t hide their tests from agents. Instead, programmers direct their agents to keep working until all tests pass." but, as we noted above, doing this results in brittle code that fails in the real world (or if you have holdout tests that aren't given to the agent, it fails the holdout tests at a very high rate; this problem cannot be solved by just adding a few more tests; it can perhaps be addressed via something like fuzzing or property-based testing, but how well that works is a topic for another post). I'm not saying these papers are bad or that there isn't something interesting to learn from these papers, but as a programmer who wants to know what techniques or tools I should use, I can't get that information from papers like the ones linked above.

    [UPDATE: Tom Adamczewski sent me a link to his paper, https://arxiv.org/pdf/2606.30182, which does handle a lot of the issues mentioned above. Relative to this post, it tries a lot more different tasks (which is great) and tries fewer languages and fewer ways of presenting tasks. One conclusion they draw in the paper that I think falls out of trying fewer languages is that language doesn't matter; even if you exclude the very obscure languges from the evals we tried here, we can observe a correlation between language popularity/usage and result quality; because Adamczewski's paper tries a lot more tasks, you can get a more complete picture by looking at this post and that paper combined than you can by looking at either in isolation.]

    [return]
show more
‘Tiktok wouldn’t open at all’: Taiwan simulates an internet blackout, with an eye on China threat
Published: 2026-08-14 03:17:54 | Created: 2026-08-14 03:39:55

Taiwan practised cutting off internet for the first time amid concerns about the safety of the undersea cables that form its ‘digital lifeline’

In Taipei, the streets swiftly empty as air raid sirens ring out and alerts ping on mobile phones. Police officers usher any stragglers into shelters, leaving the capital suddenly deserted.

In cities across Taiwan, these kinds of drills are a fact of life every year, amid the growing threat posed by China across the strait, but this time there is something different.

Continue reading...
show more
Revealed: What caused man to be partially sucked out of plane
Feed: The Latest News from the UK and Around the World | Sky News (https://feeds.skynews.com/feeds/rss/home.xml)
Published: 2026-08-14 00:25:00 | Created: 2026-08-14 03:36:54
A shattered window through which a passenger was partially sucked out during a Ryanair flight last month was caused by broken engine parts, US investigators have said.
show more
Зачем нужен ещё один инструмент обхода блокировок, если уже есть VPN
Feed: Все публикации подряд на Хабре (https://habr.com/ru/rss/articles/)
Published: 2026-08-14 03:28:14 | Created: 2026-08-14 03:29:55

День выдался долгий и по-настоящему безрадостный. Сегодня произошло то, чего мы больше всего не любим: Туннельный Котик перестал помогать многим людям, и мы, вероятно, потеряли изрядную долю доверия.Больнее всего это ударило по самым старым и лояльным пользователям - тем, кто был с нами практически с самого начала. Брешь всё ещё приходится латать, и заниматься этим придётся не один день.

Больнее всего это ударило по самым старым и лояльным пользователям - тем, кто был с нами практически с самого начала. Брешь всё ещё приходится латать, и заниматься этим придётся не один день.

Но раз уж вечер располагает к разговору - самое время ответить на вопрос, который нам регулярно задают: зачем вообще нужен ещё один проект, если есть множество готовых VPN-сервисов и решений вроде AmneziaWG? Можно в конце концов поднять свой собственный VPN, а кто не умеет - тот, в общем, сам виноват.

В целом всё это верно. За исключением того, что мы решаем совершенно другую задачу.

Читать далее
show more
US says dozens of countries helped China dodge Trump's tariffs
Published: 2026-08-14 08:39:46 | Created: 2026-08-14 03:27:56
A new US report said China had moved goods through nations with lower tariffs to dodge higher levies.
show more
BBC Verify speaks to woman who filmed ICE agent pointing gun at her
Published: 2026-08-13 20:37:14 | Created: 2026-08-14 03:25:57
BBC Verify has examined footage shared by a Virginia woman in which an ICE agent points his gun at her.
show more
Indian solar mission's new findings throw light on enduring Sun mysteries
Published: 2026-08-14 02:13:27 | Created: 2026-08-14 03:24:56
Why is the Sun's corona millions of degrees hotter than its surface? And how does it maintain its inexplicably high temperature?
show more
Clacton byelection: Farage beats Count Binface but skips result announcement – UK politics live
Published: 2026-08-14 08:21:00 | Created: 2026-08-14 03:24:55

Reform UK leader says result ‘speaks for itself’, as Essex police deny that they warned him not to attend vote count

This election is a novelty with a record 34 candidates, including Howling Laud Hope of The Official Monster Raving Loony Party and Marcus White from the Everyone is God Party.

They have pounded the streets of Clacton over the past few weeks as they seek to win the support of about 80,000 voters.

Continue reading...
show more
US says dozens of countries helped China dodge Trump's tariffs
Published: 2026-08-14 08:39:46 | Created: 2026-08-14 03:21:56
A new US report said China had moved goods through nations with lower tariffs to dodge higher levies.
show more
UK politics: Country a ‘tinderbox’, says Burnham, as he calls in military and bans disposable barbecues – as it happened
Feed: World news | The Guardian (https://www.theguardian.com/world/rss)
Published: 2026-08-14 15:59:23 | Created: 2026-08-14 03:19:56

PM also says fire and rescue services will be given more cash

This election is a novelty with a record 34 candidates, including Howling Laud Hope of The Official Monster Raving Loony Party and Marcus White from the Everyone is God Party.

They have pounded the streets of Clacton over the past few weeks as they seek to win the support of about 80,000 voters.

Continue reading...
show more
Putin can no longer claim victory in Ukraine, Nobel Peace Prize winner tells BBC
Published: 2026-08-13 04:55:58 | Created: 2026-08-14 03:17:57
In an exclusive interview with the BBC’s Steve Rosenberg, Dmitry Muratov says the Kremlin leader can only "destroy Ukraine, not conquer it", with the war now in its fifth year.
show more
8/13: CBS Evening News
Published: 2026-08-13 22:30:00 | Created: 2026-08-14 02:49:55
Over 36 million under threat of severe storms; Possible plea deal in Luigi Mangione trial.
show more
Trump imposes new drone tariffs, orders Navy to allow shipbuilding abroad
Feed: ABC News: Top Stories (https://abcnews.go.com/abcnews/topstories)
Published: 2026-08-14 02:45:15 | Created: 2026-08-14 02:47:55
Trump is slapping “ad valorem” tariffs on drones and drone parts, citing the “national security threat” posed by such devices. 
show more
Trump orders Pentagon to redesign US aircraft carrier to use steam catapults to launch jets
Published: 2026-08-14 02:03:43 | Created: 2026-08-14 02:33:55

President has suggested steam catapults every year since 2017, saying digital systems are too ‘complicated’

After nearly a decade of loudly complaining about the fact that modern US aircraft carriers use electromagnetic systems to launch fighter jets instead of steam-powered catapults, Donald Trump has directed the Pentagon to redesign a new aircraft carrier to replace the modern system with the old-fashioned catapults he prefers.

Trump signed a national security memorandum on Thursday directing the defense secretary, Pete Hegseth, and the acting navy secretary, Hung Cao, to make the change he has had his heart set on since at least 2017, when he first described his idea to Time magazine.

Continue reading...
show more
Page 355 of 1017 (50824 total items)