Kernel Recipes
2026 - Security in the LLM age
Greg Kroah-Hartman’s Kernel Recipes 2026 talk examines what LLM-based
security tools mean for the Linux kernel and open source maintainers.
Speaking from direct experience with the Mythos framework (he was
publicly named as one of the people with access), he dissects
Anthropic’s widely publicized “79 kernel bugs” report, shows how few of
those bugs were real, and argues that the security community should
neither panic nor celebrate: LLMs are fuzzy pattern matchers, not
intelligent analysts. His core message, repeated throughout, is “do not
panic” — sit down, fix the bugs, and grind through it the same way the
community did with fuzzers and Y2K. The second half of the talk is
practical advice for developers and maintainers on handling the flood of
LLM-generated reports and patches, including how to spot them, how to
push back, and why you should never upload non-public data to these
tools.
Key Points
- The kernel security team is now processing roughly 33 CVEs a day, up
from the ~50 per week Greg reported a year earlier, but their tooling
scales fine; it’s everyone else’s infrastructure that is panicking.
- Anthropic’s Mythos report claimed 79 kernel bugs. Auditing the raw
data: 24 were crashes with no substance, 14 weren’t bugs at all, 3 were
completely fabricated, 4 were already fixed by Anthropic, 11 were
publicly known and already fixed, 4 were trivial authenticated-NFS
issues, and after deduplication only 20 fixes were needed — about ten of
them real, amounting to roughly one hour of normal kernel
development.
- Of the “real” bugs, most assumed threats the kernel explicitly does
not consider security issues: malicious handcrafted filesystem images,
root-only packet injection, no-MMU corner cases, enterprise-only SCTP,
and minor IPv6 issues.
- LLMs are “fuzzy pattern matching” over code, not intelligence — the
same idea Sasha Levin and Julia Lawall pioneered years ago with
Coccinelle-style semantic patching. They work well on code but cannot
count and are sycophantic: ask for a bug and they’ll scrape the web for
bugs other people already found and fixed.
- The real threat is the exploitation window: time from patch
availability to in-the-wild exploitation is now around negative seven
days, and LLMs can chain many minor issues together to get access. The
bill is coming due for years of unfixed deployed software — banks are
finally planning to update their systems.
- The response is the same as with fuzzers and Y2K: do the work, fix
the bugs, and the noise stops. The rsync project is proof — after
systematically fixing everything scanners found, the latest rsync comes
back clean.
- Roughly half of LLM-generated patches that look correct are flat-out
wrong, based on Greg’s research with six graduate students reviewing
large batches of them. Treat any LLM-generated patch as wrong until
proven otherwise.
- Common LLM patch tells: walls of persuasive changelog text (deleting
it and reviewing just the code works better), excessive comments for
tiny changes, pointless new booleans/flags, obsolete patterns learned
from decades-old kernel history, and the dangerous habit of changing
mutex_unlock to mutex_destroy.
- These tools leak. Anything you upload will reach someone else —
reports get scraped and re-reported, and companies have used uploaded
data for other purposes. Never upload non-public information; run local
models instead.
- The kernel community’s countermeasures: documented per-subsystem
threat models (so bots and reporters can be pointed at them), requiring
security reports to CC maintainers and the security list, demanding a
patch up front (which flushes out many non-bugs), a full-time funded
security developer at kernel.org (OpenSSF and Alpha Omega), and OpenSSF
tooling for threat-model documents.
- Maintainers should push back: ask how things were tested, ask for a
reproducer, don’t take batches of a hundred patches, and remember there
is no obligation to accept any of this. People who can’t explain their
patch didn’t do the work (“don’t be a meat puppet” for an LLM).
- Expect a rough 12–18 months, then it gets better. Code is being
deleted (unused protocols, dead drivers), the share of real bug fixes
landing in Linus’s tree is rising, and the community has been through
this before with fuzzers. When the bugs are fixed, they stay fixed.
Detailed Summary
Opening: the CVE flood
Greg opens with his annual talk for the kernel community (developers
and maintainers specifically, not the general public). A year ago he
reported the kernel producing ~50 CVEs a week; it is now 33 per day. The
tooling he, Lee Jones, and Sasha Levin built for the earlier scale has
paid off — the kernel team’s infrastructure handles the load fine, while
the wider CVE ecosystem and CNAs are the ones panicking and rebuilding.
His refrain for the talk: do not panic, and ignore the doom
marketing.
The Mythos story
In early February he got the phone call about Anthropic’s Mythos
report claiming 79 kernel bugs, and got access to the raw data. His
breakdown of the 79:
- 24 were “nothing” — a crash with no report, no substance.
- 14 weren’t bugs at all; nothing happened.
- 3 were completely made up (he avoids the word “hallucination” — the
data was simply fabricated).
- 4 were already-fixed bugs the tool scraped from public mailing lists
and re-reported.
- 4 were minor authenticated NFS issues that were fixed quietly.
- 11 were famously found and fixed in public before the report came
out.
- That leaves 26 “actual bugs”, which after the tool’s inability to
count resolved to 20 fixes.
Of those 20: seven assumed a malicious handcrafted filesystem image
(the quintessential newbie non-security bug), one assumed root could
inject packets mid-stack, two were no-MMU bugs (including an io_uring
no-MMU bug he fixed — and evidence that nobody actually runs no-MMU
systems), six were SCTP bugs relevant only to enterprise telco networks,
five more were minor IPv6 issues of minor issues (two duplicated), and
one was a GPU driver issue requiring an already-malicious local user.
The end result: about ten real fixes, roughly one hour of kernel
development at the normal pace of ~10.5 patches per hour. Marketing
inflated it to a global story.
He notes the Mythos framework itself is well built — it spun up a
no-MMU RISC-V VM, wrote a test case, and exercised io_uring — but it is
a framework around a dumb chatbot. He also mentions he was doxxed by Jim
Zemlin as the only named person with Mythos access, and that he found
and fixed more real bugs in an hour than Anthropic’s report did.
The real problem:
deployment, not discovery
The genuine danger in 2026 is not that the tools find bugs — it’s
that people don’t roll out fixes. The classic discover → disclose →
patch → deploy cycle used to give ~63 days before exploitation appeared;
it’s now around negative seven days. LLMs persist relentlessly and can
chain many minor issues into real access, which makes fixing even minor
stable-tree bugs important. But the deeper cause is years of unpatched
deployed software — “the bill is finally coming due” — and Greg is now
hearing banks say they will actually update their systems, fifteen years
after being told to.
Grind through it:
fuzzers, Y2K, and rsync
His prescription is historical: the community survived Y2K and the
fuzzer storm of six or seven years ago by sitting down and doing the
work. Fuzzers still hit the kernel daily; nobody panics anymore. When
bugs are fixed, they stay fixed. The rsync project is the showcase:
Tridge and others ground through everything the scanners found, rebuilt
the testing infrastructure, and the latest rsync now comes back clean
from every code scanner. CNCF core cloud projects scanned clean too.
These tools can go deeper than fuzzers because they’re static analysis —
the same idea as Coccinelle from decades ago — but the noise stops once
the bugs are fixed.
Greg is blunt: these companies “sucked up our data” and admit they
obliterated copyright (amusingly serving an FSF goal), while
compensating no one — so developers shouldn’t pay for the tools either.
He cites LG Research in Korea finding that only ~20% of the foundational
training data is legally usable. His attitude: they spent billions
building a fuzzy pattern matcher that finds bugs in open source — fine,
let’s use it against them and make our software better.
Handling the flood of
reports and patches
Practical guidance from his past year of research (six graduate
students reviewed large volumes of bot output):
- Assume half of correct-looking LLM patches are wrong — they don’t
fix the problem, the code path is unreachable, or there was never a bug.
Patches have even fooled Greg.
- Watch for the classic
mutex_unlock →
mutex_destroy substitution that would badly break
things.
- Delete the wall-of-text changelog and review the code only; his
interns’ instinct to do exactly that turned out to be the best
approach.
- Beware obsolete patterns: models trained on decades of kernel
history reproduce old insecure idioms (insecure CGI/Perl era code being
the analogy), and Coccinelle shreds the results.
- Push back: ask how it was tested, demand a reproducer (as Jens did
to Greg), and remember no maintainer is obligated to accept anything.
Networking maintainers’ “Christmas tree” header formatting is cited as a
human-detection test.
- The flood includes inflated vendor claims: one company’s “100 bugs”
reduced to two minor fixes and one hardening issue; another’s 100 bugs
had ~49 non-bugs and the rest not security issues.
Leaking and confidentiality
Anything uploaded to these services leaks — to other users, to the
vendor, to training. He cites the mathematicians’ experience with Claude
and the weekly stream of people angry that “their” bug was reported
publicly a day earlier after being scraped. The rule: run local models,
and never upload anything non-public. Offers to pipe the
security@kernel.org feed through someone’s cloud model are flatly
refused.
- Documented per-subsystem threat models (add yours!), which let the
security team point bots and reporters at documented assumptions; eBPF
and networking are well covered, filesystems still get weekly
handcrafted-image reports. OpenSSF offers a tool to generate
threat-model documents for user-space projects.
- Security reports now CC subsystem maintainers and the security list;
maintainers should handle them at normal pace and push discussion to the
public mailing list when appropriate.
- Reporters are asked to send a patch up front — the act of producing
a patch filters out many non-bugs, and reporters want credit
anyway.
- OpenSSF and Alpha Omega fund a full-time security developer working
with kernel.org.
- Shishko, the open-source local code-review bot, plus agents written
by Chris Mason (kernel development) and Takashi (sound subsystem), run
locally.
For developers and
maintainers
- Push back on anything that looks wrong; ignore doom marketing
(you’re not the buyer).
- Run local models on your own hardware — they’re good enough and free
to use.
- It’s fine to delete code: one intern “fixed” a driver by deleting
3,000 lines nobody used.
- It’s okay to accept genuinely helpful edge patches (e.g. people
getting audio working on their laptops with LLM help), but core code is
a different story.
- Staging and mentorship programs have banned LLM-generated patches
unless the contributor has the hardware and test evidence; the LFX
mentorship program requires students to show their testing.
- Trust is the currency: prove you’re human, respond to email, send
one or two patches rather than a hundred (Greg’s own Kubernetes patches
were rejected as bot-like), and go to conferences. “Just act human.” A
contributor who can’t explain their patch doesn’t deserve authorship —
“don’t be a meat puppet.”
Q&A highlights
- Using a second LLM to parse another LLM’s 4,000-character
single-line reports is wasteful; keep insisting on patches from the
analysis, not reports. Some Linux Foundation groups are helping triage
these walls of text.
- On AI accountability: one project introduced confidence scores and
inspectable chat histories with good results, but Greg doubts that
scales to 5,000+ kernel developers; getting people to use the
Assisted-by tag is already hard enough.
- Bug-fix percentages landing in Linus’s tree are creeping up (9 → ~11
patches/hour); the bug-introduction curve is flat, with 5.15 an outlier
due to the in-kernel SMB server.
- On productivity: “I do not feel more productive. I have to process
more patches.” Measuring that is management’s job, not an open source
project’s.
- Maintainers handling security reports alone for the first time
should ask the security team for help — that’s what they’re there for,
and load is actually dropping as reports spread across maintainers.
- His prediction: a rough 12–18 months, then the community comes out
the other side with better code, just as it did with fuzzers.