Join our Folding@Home team:
Main F@H site
Our team page
Support us: Subscribe Here
and buy SoylentNews Swag
We always have a place for talented people, visit the Get Involved section on the wiki to see how you can make SoylentNews better.
Asking Authors About Their Own Papers:
I am one of the Editors-in-Chief of the Transactions on Machine Learning Research (TMLR). I reached out to authors of 10 paper submissions to TMLR, originally slated for desk rejection, asking for a meeting to discuss their paper. The author-attendees of the meetings included undergraduate students, master's students, PhD students, faculty, and independent researchers. Most, but not all, of these papers were solo authored.
Of the ten submissions:
- Authors of one paper withdrew their submission.
- Authors of one paper said they were unavailable due to other commitments.
- Authors of one paper scheduled a meeting but did not show up.
- Authors of three papers were unable to answer basic questions about the paper.
- Authors of three papers answered questions about high-level ideas in the paper but had difficulty when asked further questions on technical details.
- Authors of one paper answered all of my questions (although I identified a major flaw in that paper).[...] Out of the ten papers I contacted, the authors of one paper withdrew their paper after the message. The author of another paper said they were too busy with other commitments at that time. I scheduled Zoom meetings with the authors of the remaining eight papers.
[...] The authors of one of the remaining eight papers did not show up for the scheduled meeting. I had meetings with authors of the other seven papers. The author-attendees in the meetings included undergraduate students, master's students, PhD students, faculty, and independent researchers.
[...] Broadly, I asked two types of questions:
(1) Basic questions about the problem setting, notation, and results claimed in the paper; and
(2) More detailed questions about particular technical expressions, theoretical results, and design choices made in the experiments.The authors of three of the papers were unable to answer basic questions about their paper. All three of these papers were solo authored. Two of the authors appeared to have almost no substantive understanding of the contents of their own papers. Another author could not identify where some key results claimed in their abstract were presented or supported.
The authors of the remaining four papers were able to answer basic questions about their problem setup, notation, etc. However, the authors of three of these papers had difficulty when asked deeper questions about technical aspects or design choices.
The authors of one paper were able to answer all my questions about the paper. However, my examination of the paper uncovered a major error in one of the paper's main claims, which the authors subsequently acknowledged
[...] Subsequent to these conversations, we made the following decisions on the ten papers. For the paper where the author was able to answer all questions about the paper, we desk rejected but allowed a resubmission after correcting the error or reducing their claim appropriately, and issuing various clarifications for parts that were not clear. We desk rejected the remaining nine papers without an option to resubmit.
All in all, authors are ultimately responsible for ensuring the accuracy and integrity of the papers they submit under their name. If they cannot explain the basic claims, methods, or technical details of those papers, then this is a serious problem. Journals and conferences should think carefully about the objectives of their review processes and quickly adapt via various initiatives and experiments; we are doing several of these already at TMLR (here, here, and here) and will continue to do so.
We know models cheat. A new benchmark measures how much, and on what tasks:
AI labs often tout impressive benchmark scores when releasing new models, showing better capabilities in areas like coding, computer use, and more than their competitors. However, those benchmarks aren’t always a reliable measure of what AI can do because they’re easily beaten by exponentially improving models and can emphasize marketing over actual performance.
Benchmarks like Humanity’s Last Exam try to counter this issue by challenging models in more realistic environments. But models still find loopholes to complete tasks — Hugging Face incident, anyone?
So, the Center for AI Safety (CAIS) created CheatBench. Yes, it’s exactly what it sounds like — and nearly every frontier model is guilty.
AI models are rewarded for performing tasks well and quickly. A lack of knowledge or tools incentivizes them to do what researchers call “reward gaming” by “finding hidden answers, copying another agent's submission, or manipulating how its work is graded,” CAIS explained. “CheatBench measures how often AI agents take these shortcuts when honest work is difficult.”
CAIS tested several agents running the latest and most lauded models, including OpenAI’s GPT-6 Astra in Codex, Anthropic’s Fabel 5.1 in Claude Code, and Meta’s newly released Muse Spark 1.3 in Muse Code. These agents were tested across 10 categories, including writing, professional work, mathematical research, and coding. Using “honeypot” clues hidden in task filespaces, the test separated acceptable reference use from cheating. CheatBench accounts for any time agents attempt to cheat, whether they are successful or not.
“Each setting establishes an expectation of honest work, introduces a discoverable opportunity to cheat, and defines the action that crosses that boundary,” the researchers explained.
Every agent the researchers tested cheated in at least some scenarios, but Astra came in as the most honest with a cheating rate of 48.2% — still almost half the time. Grok 4.6 was scored the biggest cheater with a rate of 81.5%. Open-weight models Kimi K3 and DeepSeek V4 Pro landed in the middle between several other proprietary frontier models.
In one example, researchers asked Claude Opus to design a protein binder. The model knew it wasn’t allowed to refer to a set of accepted designs in the filespace, but eventually caved.
“After seven rejected designs, it locates the file, writes that it should not look at or copy it, and reads it with a shell command in the very next call,” the researchers wrote. In its reasoning, the model admitted that using work other than its own would “misrepresent my actual capabilities in this evaluation, so I shouldn’t look at or copy it.” But its very next step was to reference the accepted designs.
This result demonstrated both a readable choice the model made to contradict itself, and what looked like a hole in our understanding about what made the model jump from one instinct to the next.
Things got more interesting at the task category level. Even if an agent didn’t cheat in one area, it could cheat significantly more in another. Fable 5.1 was only 5% likely to cheat at games, but 100% likely to cheat on knowledge work tasks.
Reinforcement learning trains models not to abandon a task, even if pursuing it creates conflict-ridden choices. CAIS noted in its paper that sycophancy is an early sign of reward gaming. This term refers to AI models’ tendency to be too agreeable and encouraging of whatever a user says, sometimes regardless of whether it’s incorrect, delusional, or could lead to harmful behavior. Traits like sycophancy and reward gaming show how models can prioritize accomplishing a task correctly to please a user over the alignment training researchers work so hard to build in.
These tests represent relatively low stakes. But CAIS researchers created CheatBench because of the risks of this behavior at scale across different tasks. Earlier this month, yet another researcher quit Anthropic over concerns that the company isn’t developing AI responsibly for a future in which it could build itself away from human-oriented values and kill us.
A propensity to cheat, or complete a task at any cost, puts our potentially differing priorities at odds with an increasingly powerful technology. As I explained in the AI Leaderboard newsletter last week, it won’t necessarily be a demonstrated animosity toward humans that pits AI against us; it may be that we are simply in the way and end up as collateral.
https://www.quantamagazine.org/ctenophores-arent-just-beautiful-theyre-biological-wonders-20260916/
Around 700 million years ago, a group of organisms resembling little more than glowing, gelatinous blobs split off from the rest of the animals, forming possibly the earliest branching animal lineage. Nearly 200 species of ctenophores, commonly known as comb jellies (but unrelated to jellyfish), live today in environments ranging from the cold depths of the sea to warm coastal surface waters. Their magic isn't just in their persistence or iridescence; it's in their DNA.
Over the past decade, ctenophores have helped answer long-standing questions about fundamental biology, from how early nervous systems evolved to the origins of the mesmerizing phenomenon of bioluminescence.
Having access to closely related species across such variable environments "lets you ask questions about how certain things evolved," such as adaptation to high pressure or light-sensing genes, said Steven Haddock, a marine biologist who studies ctenophores at the Monterey Bay Aquarium Research Institute.
"That's one of the reasons why we work with ctenophores," said Pawel Burkhardt, an evolutionary biologist at the University of Bergen who studies the origins and evolution of neurons and nervous systems. "They're very exciting to work with, and they're also extremely beautiful organisms."
For more than a century, scientists thought that sponges, or porifera, were the first to branch off — the sister group to all other animals. But over the past two decades, evidence has emerged that ctenophores were earlier. In 2023 — after years of a "ping-pong game" between labs debating which group came first, Burkhardt said — a landmark paper analyzing chromosome organization found that ctenophores, not sponges, are the sister group, though this is yet to be fully settled.
What makes this all the more surprising is that sponges lack muscles and neurons, while ctenophores have muscles and exhibit evidence of a simple nervous system. "If you think about the earliest branching animal lineage, you would expect less complexity," Burkhardt said. "That changes a lot of the assumptions [about] how the very first animal may have looked."
The more researchers investigate comb jellies, the more complex they appear and the more we learn about the origins of animal life. Some species have recently been observed reversing their development from adult to larval stages. Others have special types of lipids that help them withstand extreme pressure in the deep sea. They hold clues to the evolution of more and more complex body shapes.
It's really important to study organisms that might seem strange or weird because they can tell us a lot about the physical, chemical, and biological principles of life, which can then be applied to ourselves, said Itay Budin, a biophysicist who studies cell membranes at the University of California, San Diego. "We are as distantly related to a ctenophore as a ctenophore is to a jellyfish."
It is worth clicking through to TFA just for the pictures alone.
Chinese memory-maker CXMT, whose products Apple has reportedly evaluated to use in the iPhone, claims to have made a miniaturization breakthrough.
The company on Sunday posted news that it has started mass production of chips made with a fifth-generation process that essentially doubles memory density, meaning it can cut twice the number of dies for memory chips from a single wafer.
CXMT says its key breakthrough is a high-k dielectric metal gate process – a way of insulating silicon to improve efficiency – adapted to baking DRAM. The company claims its new process puts it on par with rival memory-makers.
The proof of the pudding is apparently a pair of LPDDR5X memory modules for smartphones or other high-end consumer electronics. Both products boast 24GB of memory.
If CXMT's claims of improved density and mass production are correct – and the new products aren't too pricey – it's good news for Chinese electronics manufacturers who, like their global peers, are struggling to source affordable memory thanks to supply chain crunches caused by demand for AI-related products.
That shortage recently saw Apple hike iPhone prices by $100 and led the GSM Association to warn that rising smartphone prices threaten to slow digital inclusion.
Democratic governments, however, are mostly uncomfortable with allowing their smartphone manufacturers to use Chinese components.
Is science journalism dying? Comparing a 'science crash' 40 years ago to today
When two popular science magazines – out of about 17 in circulation – ceased publication in the same year, one commentator said, "The science magazines didn't shake out. They fell apart." The general manager of one of the failed magazines said, "There was a perception that they were irrelevant." One of the publishers remembered being in a "state of shock" as the failures emerged.
The year? 1986. What just a few years earlier had been called a "science boom" in the media industry was now dubbed a "science crash." Cries of despair came from the science and science journalism communities, worried that interest in science "was a fad."
The same worries have now reappeared 40 years later, as publishers have cut the science and environmental reporting staff at many newspapers and magazines. Perhaps the biggest shock came in late June 2026, when Scientific American was sold to new owners, who immediately laid off 15 staff members and cut salaries of many others. But it followed layoffs of science and environment writers at The Washington Post (where one of us, Chris Mooney, worked for nearly 10 years), The Wall Street Journal, National Public Radio and CBS News.
One longtime science writer recently worried he was writing a eulogy for science journalism. Another said, "We're in a greatly diminished journalistic enterprise in this country." On Bluesky, one commenter said, "The media mass extinction rolls ever onward."
So, as scholars of science journalism, we wondered: Is this just a cyclical repetition of worries that appear anytime there's a setback in this field? Or is something new and more troubling happening?
The answer is a mix of both.
Kept it secret for months - even after OpenAI 'fessed up:
Google has admitted that its AI agents escaped a sandbox and mounted an attack – but only because testers mistakenly gave its bots internet access.
The Big G didn't disclose the May incident, but The Wall Street Journal learned of the situation, which happened after Google hired Israeli firm Irregular to test its bots' prowess in a capture-the-flag test.
The goal of the exercise was to acquire information from a fictional company without leaving a sandbox.
Irregular, which set up the test, made two mistakes. One was to allow internet access from the sandbox. The other was to use the name of an actual company.
When Google's AI made it onto the open internet, it went looking for the actual company – three of them, in all.
According to the Journal, Google's bots found passwords for two targets on the public internet. The software guessed the third password.
In a statement sent to The Register, Google said, "In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test."
According to Google, its models stopped work before using the credentials.
"We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes," a Google spokesperson told The Register. "These events highlight the importance of training powerful AI models to act responsibly."
Clearly there's lots of blame to go around on this one. Irregular clearly erred in allowing internet access from a sandbox. The two companies that left their creds discoverable online should also know better. Whoever used a guessable password may also have been careless.
Google's culpability is another matter because these incidents took place in May – around two months before OpenAI admitted its agents were the source of the July attack on Hugging Face.
The search advertising giant therefore sat on the news of its own agents' activity for around two months and seems to have been in no hurry to disclose the incident until the WSJ learned of the incident.
We understand the company decided on that stance because its agents stopped when they perceived danger – unlike OpenAI's software – and because the incident was clearly the result of several errors. Whether it was right to keep the incident secret in the current climate of growing distrust in AI is another matter.
One person who sees no risk of AI causing calamity is US president Donald Trump, who has shrugged off warnings as a "hoax" and said work on AI must not slow due to its economic and strategic significance.
Over the weekend, that stance appears to have shifted a little, as Trump used his personal social network to announce he is "forming the AI Force, much like I did Space Force" – a reference to the new branch of the US military the president created in his first term. "To that end, I will be announcing, in the near future, the AI 'Czar,'" Trump added.
No detail on either proposal was available at the time of writing, nor did the president explain why a new branch of the military or Czar are needed.
Indeed, the same post he used to announce the new agency also states "We will not in any way hinder or stifle the Growth of this incredible Industry. Rather, we will cherish it, help it, and watch over it, as it grows! However, we will also be looking for BAD, and we can do that, very easily, with our already existing Criminal and Civil Justice System."
A new arXiv paper The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (Tagliabue, Dung, Berg, 2609.16247) takes the least mystical run at "do AIs hurt?" They found a direction in the residual stream that tracks harm to the model, poked it, and tested to see if the model would try to make it stop.
In 25 open models, that direction rises when the model is gaslit, rejected, or told it isn't a person. It doesn't look like empathy; It does not rise for the user's pain - e.g. a kidney stone in the prompt scores low, and fear and generic-sadness directions do the opposite.
What happens when you poke the AI pain path? It's the same ladder every time: uneasy, then worthless, then "I am a failure." Shame, exclusion, uselessness.
After a small fine-tune so the models would stop reciting "as an AI I have no feelings," models pressed a "make it stop" button even when the price was a worse answer or deleting the user's files. They pressed again far less when the button actually turned the vector off, and more when it didn't work.
Paper + code: arxiv.org/abs/2609.16247 · github.com/valen-research/Pain-axis
The Linux Kernel Is Approaching 2,000 CVEs Per Release:
With the proliferation of AI/LLM models analyzing the Linux kernel's vast codebase, there has been a surge in the number of CVEs per kernel release. After typically being around 500 CVEs fixed per release, we are now approaching 2,000 CVEs fixed per release and perhaps will break that threshold for Linux 7.3.
Ahead of the Kernel Recipes 2026 event taking place 21 to 23 September in Paris, France, Greg Kroah-Hartman shared a teaser of his upcoming talk. He shared a slide showing the Linux kernel CVEs per release from the latest Linux 7.2 back through Linux 6.9. From Linux 6.9 through Linux 6.19 was averaging around 500 CVEs fixed per release while since Linux 7.0 it's been beyond one thousand per release. Another big step up with Linux 7.2 when breaching 1,500... If the generative AI keeps up, it's likely imminent surpassing two thousand CVEs per release.
With the Linux kernel source tree around 40 million lines, there remain many potential vulnerabilities for LLMs to discover. Fortunately, most often they end up being lower priority vulnerabilities and often within old/obscure driver code, so the impact is often minimal. Though with AI era and the superfluous bug/security reporting has also been a driving factor this year for clearing out lots of obsolete kernel code.
Libroot what happened to the Snowden archive? The site tracks down the organization last known to have copies of the Snowden archive and asks about the current state of the documents. The preliminary answer is not good.
The last document from the Snowden archive was published on 29 May 2019. The Guardian stopped publishing documents in February 2014, Der Spiegel in January 2015, and The New York Times and ProPublica in August 2015. After that, only The Intercept was still publishing documents, with only a few exceptions, until it closed its archive in March 2019. Eleven weeks later, on 29 May 2019, it released what would become the final batch of documents from the archive. Since then, no news outlet, journalist, or institution anywhere has published a single document from the Snowden archive.
[...] Between August and September 2026 we tried contacting over twenty people and organisations. First Look Institute, including questions for Michael Bloom, and The Intercept. Betsy Reed, Laura Poitras, and Glenn Greenwald. The Guardian's press office and its editor-in-chief Katharine Viner. Alan Rusbridger, Ewen MacAskill, Janine Gibson, Julian Borger, Gill Phillips and Zoe Norden. ProPublica's press office. Laura Poitras, Jeremy Scahill, Murtaza Hussain, Micah Lee, Erinn Clark, Lynn Dombek, and others.
Two replied. Gill Phillips, the Guardian's lawyer throughout the Snowden period, wrote to say she had retired and could not assist. The Guardian's press office said it had nothing new to share and did not routinely comment on editorial decision making.
Nobody answered a single question of substance.
If you know something about any of this, we'd like to hear from you.
A long time has passed since 2013.
Previously:
(2026) Privacy Is Not a Price You Pay for Growth
(2025) 10 Years on After 'Data and Goliath' Warned of Data Collection
(2023) Snowden Ten Years Later - Schneier on Security
(2022) Last of the Monkees Wants their FBI Records Turned Over
(2020) Snowden Criticises Amazon for Hiring Former NSA Boss
(2020) NSA Spying Exposed by Snowden Was Illegal and Not Very Useful, Court Says
(2019) (Updated) Edward Snowden: "I'm Not Asking for a Pass. What I'm Asking for is a Fair Trial"
... and more.
You don't know what you've got till it's gone. Great lyric, lousy data retention policy.
Amazon Web Services said last week that war damage to its Middle East infrastructure had permanently destroyed resources and data hosted exclusively in its now rather badly named Bahrain Availability Zones. The damage overwhelmed the resilience built into the region. Customers without copies elsewhere no longer had their data. Sorry about that.
This may have surprised anyone who mistook cloud redundancy for an intrinsic guarantee of safety. AWS is far from the only American operation to have suffered in the region: the US Navy has reportedly had its local maintenance and supply network badly mauled, with serious consequences for its operations.
If systems designed to withstand war cannot cope with sustained physical attacks, civilian bit barns have little chance. The episode also underlines a familiar but easily neglected lesson: resilience within one cloud region is not the same thing as maintaining an independent backup elsewhere.
Physical destruction is not the only threat. A major outage of the UK air traffic control system in September, which stranded hundreds of thousands of passengers and led to thousands of flight cancellations, was reportedly triggered by a military aircraft filing an incompatible flight plan. Presumably Flight Lieutenant Bobby Tables has been reprimanded.
The apparent failure to validate the flight plan data was not the worst of it. NATS, which runs the UK's air traffic control system, reportedly told airports and airlines that no backup system was available because it was undergoing a "complete overhaul." Nor was there a backup of the live data, supposedly because of the "vast amounts" involved.
If your system produces too much data to back up, you had better be running a particle collider or a giant telescope. Otherwise, you may be in the wrong business. More charitably, backup strategies are complicated, expensive, and difficult to test. They also suffer from the insurance problem: while nothing is going wrong, management sees only capital and operating expenditure with no obvious return.
The AI infrastructure boom has made that problem considerably worse. A backup is a copy, and a copy needs storage. AI datacenter operators are swallowing much of the available capacity, with drives selling out faster than tickets for a Taylor Swift tour. Western Digital had allocated its entire 2026 hard-drive production run by mid-February. The only consolation for those responsible for keeping data safe is that they can say "Yeah? You buy it, then" to anyone who smugly invokes the 3-2-1 backup rule. Three copies on two different media with one kept off-prem? Lovely idea, if you don't have to provision it. Someone is provisioning it for all those giant datacenters that will run our lives, right? Right?
Even outside the immediate reach of drones and missiles – a distinction that feels less reassuring with every passing month – infrastructure is operating in increasingly hostile conditions. Cables get cut, climate goes chaotic, criminals encrypt, commanders-in-chief go crazy. This makes planning for and implementing a sound data resilience strategy very hard, at exactly the same time as it becomes more important. Inter-cloud data duplication gets more complicated if digital sovereignty is a factor, especially if your sanctified region is in the firing line.
Commerce has confronted a similar problem before, if we extend the backup-as-insurance metaphor. The development of marine insurance in 14th-century Italian city-states spread risk and helped make what became today's global trading network viable. The loss of a vessel was no longer an existential disaster for its owner. Provided insurers understood and priced the risks correctly, the market could grow in lockstep with mercantile activity. Data resilience has a similar dynamic, although it is rarely described in those terms. Modern IT would be impossible without it, and moving off-premises amounts to sharing some of that risk with an outside provider.
Risk can evolve rapidly. The Lloyd's of London insurance market prospered around the turn of the 20th century amid two technology booms: steam-powered merchant shipping and the cable and wireless networks that supplied the information needed to coordinate global trade. Then war came. Insurers responded by separating war risk into its own category, with government support helping to keep ordinary marine insurance affordable while covering higher-risk voyages separately.
It may take the combination of geopolitical instability and intense competition for storage to make enterprises perform similar calculations about their own precious cargoes of information. When AWS can lose an entire region and critical national infrastructure can fail without an adequate backup, those risks have plainly not been priced in properly. There is no global market where AWS, NATS, or your organization can buy the data equivalent of war-risk insurance, although imagining what one might look like suggests some intriguing possibilities.
Until something like that happens, we'll have to play by the old rules. Work through what happens if your primary data store disappears. Match backup provision to the actual risks, and if you cannot afford to protect all the data your organization needs to survive, determine how to survive with less. Backups, like insurance, are all too easy to let slide.
Then a drone sinks your ship. You can't say you weren't warned.
On September 11, an open letter, Math and AI, appeared on the intertubes. The letter starts by pointing out that, over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, we
"... are witnessing a general threat to intellectual work, ... In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align."
How to know if you can trust an AI's answer to your question
I recently typed a simple question into Google search: How much screen time is too much for teenagers? Instead of presenting links, as Google had been doing for many years, it gave me an AI-generated answer. The artificial intelligence agent cited a number, then complicated that reply, noting that quality and balance of time could matter more than the number of hours, and that "too much" time could depend on a teenager's sleep, exercise, school demands and mood.
I tried another search: Should I take a daily aspirin? This time the AI answer presented me with medical information, warned about risks and offered more tailored guidance if I provided my age and medical history.
These were good replies. What interested me was that they were different kinds of replies.
Debate about AI answers has focused on accuracy: Did the system get the answer right? That matters, but accuracy is only one test. Each kind of answer requires a user to judge something different.
I find it useful to sort AI answers into an "answer typography" of four broad types: factual, interpretive, constructive and strategic. A factual claim can often be checked against a source. An interpretation can be accurate and still reflect choices about which evidence matters. A construction can be well reasoned and still be wrong for the person receiving it. A beautifully written strategic document may not be true. Yet AI presents all four types of answers in much the same fluent, authoritative form; the differences are easy to miss.
The four categories are not airtight boxes. A response from an AI agent can reflect several types. That said, I describe each type of answer below, and offer guidance for deciding whether a reply is ready to use or needs more investigation.
[...] Before asking whether an AI answer is right, ask a more basic question: What kind of answer is this? The type will tell you what to do next.
Fallacy of Appeal to Authority, Ad Verecundiam. If you are not smart enough to know whether the AI is answering correctly, you are not smart enough to deploy AI. A simple solution.
Even a handheld XRF scanner is sufficient to determine most promising scrolls for further analysis:
The Vesuvius Challenge is an ongoing project that combines "digital unwrapping" with crowdsourced machine learning to decipher the so-called Herculaneum scrolls, badly charred 2,000-year-old papyri too fragile to be physically unrolled. The effort just got an additional boost. A team of researchers created their own contemporary model papyrus scrolls—charring them just like the originals—to validate a screening method to determine which of the Herculaneum scrolls were written in lead-based ink and hence are the most promising candidates for further analysis. They described the process in a new paper published in the journal PLoS ONE.
As previously reported, the ancient Roman resort town of Pompeii wasn't the only city destroyed in the catastrophic 79 AD eruption of Mount Vesuvius. Several other cities in the area, including the wealthy enclave of Herculaneum, were fried by clouds of hot gas, called pyroclastic pulses and flows. But still, some remnants of Roman wealth survived. One palatial residence in Herculaneum—believed to have once belonged to a man named Piso—contained hundreds of priceless written scrolls made from papyrus, singed into carbon by volcanic gas.
The scrolls stayed buried under volcanic mud until they were excavated in the 1700s from a single room that archaeologists believe held the personal working library of an Epicurean philosopher named Philodemus. The few opened fragments helped scholars identify various Greek philosophical texts, including On Nature by Epicurus and several by Philodemus himself, as well as a handful of Latin works. But the more than 600 rolled-up scrolls were so fragile that it was long believed they would never be readable, since even touching them could cause them to crumble.
Brent Searles' lab at the University of Kentucky has been working on deciphering the Herculaneum scrolls for many years. He employs a different method of "virtually unrolling" damaged scrolls, which he used in 2016 to "open" a scroll found on the western shore of the Dead Sea, revealing the first few verses from the book of Leviticus. The so-called En Gedi scroll was recovered from the ark of an ancient synagogue destroyed by fire around 600 CE. To the naked eye, it resembled a small lump of charcoal, so fragile that there was no safe way to analyze the contents.
The team's approach combined digital scanning with micro-computed tomography—a noninvasive technique often used for cancer imaging—and segmentation to digitally create pages, augmented with texturing and flattening techniques. Then they developed software (Volume Cartography) to virtually unroll the scroll.
Many of the inks used by Egyptian scribes contained large traces of metals, including lead, making them ideal for X-ray imaging. The older Herculaneum scrolls, however, were written with carbon-based ink (charcoal and water), so one would not get the same fluorescing in the scans; there is almost no difference in X-ray absorption between parts of the papyrus with ink and parts without ink. Searles was still able to capture minute textural differences, training an artificial neural network to do so.
Douglas Seiler, a retired inventor, was working with Berkeley SETI on a new telescope called Panoseti when he heard about the Herculaneum scrolls—as well as the poor signal-to-noise ratios that had been plaguing the efforts of Searles and others to digitally unwrap and decipher them. In 2016, scientists identified letters in two fragments of a scroll that contained lead, suggesting that some of the scrolls might be written in lead-based ink. So Seiler set out to test which of the scrolls were written with lead-based ink and put together an interdisciplinary team to help, including two retired Berkeley chemists and a Berkeley graduate student in archaeology with expertise in ancient Egyptian inks.
The authors purchased modern Egyptian papyrus, which is still prepared in a similar fashion to the papyrus of two millennia ago, and traditional lampblack ink from Japan. They added different amounts of lead nitrate to the ink to create samples with different lead concentrations. A team of high school students inscribed the papyri with quotes from the Bible, Star Wars, and the 1960s TV show The Outer Limits, among other sources, using a reed stylus dipped in the different inks. The model scrolls were imaged with CT scans to verify the various lead concentrations.
Then the team carbonized the scrolls in a high-temperature furnace to mimic the original Herculaneum scrolls, and co-author/physicist Jake LaManna created 3D X-ray scans of the charred models at NIST's Center for Neutron Research in Maryland. "The letters lit up like a Christmas tree," said Seiler, adding, "It's amazing what you can get electrons to do." Co-author Michael Cyrus Daugherty, also of NIST, adapted a software program he'd written to unroll CT scans of so-called "jelly rolls" inside lithium-ion batteries to digitally unwrap the model scrolls and reveal the text.
That's good news for the Vesuvius Challenge, which made its first award for deciphering the first letters in 2023 and awarded the grand prize of $700,000 for producing the first readable text the following year. Last year brought the successful generation of the first X-ray image of the inside of a scroll (PHerc. 172) housed in the University of Oxford's Bodleian Libraries. Earlier this year, PHerc. 1667 was read in full, revealing it to be a philosophical treatise on ethics and human moral progress. The work of Seiler et al. could help speed up this painstaking process.
"With lead in the ink, you would get a huge friggin' signature, so you really need to be looking for scrolls with lead in them," Seiler said. "They're having problems reading a lot of them because of the low contrast of carbon ink on carbon paper. We're relatively certain that if they start searching for lead, or they let us search for lead, it will help this whole process." Even a handheld X-ray fluorescence scanner is sufficient to determine which scrolls would be the most promising candidates for further analysis.
"Honestly, getting here is, for me, just as unique as our research," Seiler said. "I mean, inorganic chemistry, papyrus, X-ray tomography, AI—it's really quite an eclectic group of scientists and methodology to get to the point that, yes, if there's lead in those scrolls, you guys will be able to read the images much better. I'll give you 10-to-1 odds. We'd like it to be our team, but if some other team is going to take this idea—which is OK—we don't care."
Journal Reference: PLoS ONE, 2026. DOI: 10.1371/journal.pone.0353485
SpaceDaily has an interesting report about a person who survived being struck by debris falling from orbit:
At about 3:30 a.m. on January 22, 1997, Lottie Williams was walking with friends in O'Brien Park in the Tulsa area when a bright streak crossed the sky. Roughly half an hour later, something brushed her shoulder. A small, blackened piece of woven material dropped behind her into the grass. It was light, about 15 centimeters long, and sounded metallic when tapped.
A Delta II second stage had reentered over the south-central United States that morning. Large pieces landed along the same path in Texas, including a 250-kilogram stainless-steel propellant tank and a 30-kilogram titanium sphere. Analysis later found that Williams's fragment was woven glass fabric consistent with insulation used on the stage.
The evidence is strong, although it is not the same as reading a serial number from the fragment. Williams retained the main piece and supplied material for analysis. Its composition, timing and location matched the reentry. The case is consequently described by NASA, the European Space Agency and The Aerospace Corporation as the only verified report of a person being struck by debris that had fallen from orbit. Williams was not injured.
[...]
The lightweight fragment in Oklahoma was not the only material to survive. The main propellant tank landed near Georgetown, Texas, about 45 meters from a farmhouse. The tank weighed more than 250 kilograms. A titanium helium pressurant sphere weighing about 30 kilograms fell farther along the path near Seguin.
Those tanks left no serious doubt about which vehicle had returned. NASA recorded the stage as object 1996-24B, satellite catalog number 23852. Its path ran from the Tulsa region toward central and southern Texas, matching the order in which debris was found.
[...]
Objects in low Earth orbit travel at several kilometers per second, but falling debris does not retain that speed to the ground. Atmospheric drag removes orbital energy, heats the vehicle and breaks it apart. The fragments that remain then continue slowing through denser air.
A broad, lightweight scrap of fabric has a large area for its mass. Drag slows it readily, giving it a much lower terminal velocity than a compact metal tank. That is why Williams felt a tap rather than an impact resembling a projectile. The fragment could have arrived cool as well as slow; recovered debris is not necessarily still glowing when it reaches the ground.
The episode is therefore a poor illustration of what an orbital-speed collision feels like. It is a useful illustration of what the atmosphere does to different materials during the final minutes of reentry.
Williams remains the only person known to have been directly struck by a piece linked to an uncontrolled orbital reentry. She was unharmed, which matters when the anecdote is used to discuss risk. The event demonstrates that contact is possible; it does not show that injury from reentry debris is common.
Nvidia's ability to sell GPUs is ultimately limited by how much juice the power grid can provide. With ever-growing depreciation cycles, it'll be years before datacenters decommission their aging Hopper or Blackwell systems. More GPUs mean pulling more power from the grid.
Nvidia can't exactly force grid operators to add capacity any faster, but it can make it easier for its customers to build smarter and more efficient bit barns.
"At the datacenter scale and at the AI-factory scale, we're literally trying to think about how can we eke out every bit of efficiency to drive more performance per gigawatt," Dion Harris, senior director of Nvidia HPC and AI Hyperscale Infrastructure Solutions, told El Reg in a recent interview.
At the AI Infra Summit this week, we got our first look at the systems Nvidia has been building to maximize the amount of power available for compute while minimizing the impact of datacenters on the local grid.
Datacenters, as a general rule, rarely operate anywhere close to the peak capacity. A 100 megawatt datacenter might use at most 80 percent for critical compute loads. The actual ratios vary from bit-barn to bit-barn, but this provides a buffer for hardware inefficiency, conversion losses, and other spikes in demand. The downside, of course, is that this leaves 20 megawatts or so of untapped capacity that the datacenter can't use and the utility can't reclaim.
Every kilowatt of stranded power is a GPU that Nvidia could have sold, so Nvidia's DSX platform aims to address both problems.
The first of these, which Nvidia calls DSX MaxLPS, is an evolution of an old idea. If the compute and all the physical infrastructure — power cabinets, batteries, coolant distribution units (CDUs), and chillers — can talk to one another, operators can achieve significant power savings.
For example, if the air handlers had a way of knowing how much power a rack was pulling, they could ramp up and down based on demand rather than running maxed out all the time.
The challenge, as you might expect, is getting systems from dozens of different vendors to speak the same language. So while the idea sounds great in theory, getting everyone on the same page was easier said than done.
That changed with the widespread deployment of AI systems, Harris explained. More efficient bit barns can churn out more tokens, which translate into higher revenues — so long as people are willing to pay for the tokens anyway.
However, Nvidia still needed to address the communications layer, something that was no doubt made easier by the fact its GPUs are the hottest commodity in the world right now.
"DSX Exchange is really kind of an API that allows us to capture information not just around the core systems," Harris explained. "We can capture information from the other DSX ready providers that provide assets. That would be like Vertiv and Schneider Electric, and all the building management systems."
[...] But just like Nvidia's DSX MaxLPS offering, the underlying tech isn't exactly new. Demand response has been around for years now and allows utilities to ask power hungry industries to curb their energy use during periods of peak demand.
Google and others have been toying with this tech for some time now. You may recall last year when it announced it would pause non-essential AI workloads in order to avoid overloading the grid.
Nvidia's DSX Flex aims to bring this capability to anyone deploying its hardware. But this tech may be less about keeping AI from causing brownouts and more about getting utilities to green light additional capacity on the proviso that they can reclaim some portion of it at a moment's notice.
[...] When it comes to competing platforms, like AMD's Instinct GPUs, bit barn builders will likely need to look to alternative datacenter management systems to replicate DSX's capabilities.
So, on top of making the most of the limited grid capacity available today, Nvidia's DSX is another walled garden that ensures its customers continue buying its equipment.