Claude Code is a large codebase and uses tons of Node’s features either directly or indirectly through dependencies. fetch(), node:http, node:tls, node:os, node:net, node:fs, AbortSignal, node:child_process, node:tty, node:process, node:http2, etc.
Bun’s Rust rewrite shipped in Claude Code over a month ago and barely anyone noticed. Claude Code is widely used. The Rust rewrite is going well overall.
In the Bun v1.4 video, I promised a certain number of newly passing Node.js tests were added to force us to improve compatibility, and that number is not true yet. The release is delayed until it is true. The PRs to make it true are up but not merged yet. Most likely next Tuesday we’ll do the release of 1.4.
Take as long as you need to ensure software quality. A month without a release isn't a big deal and whomever needs a specific feature can offer to contribute or build themselves.
Node has 4-6 weeks without a meaningful release (other than security stuff) pretty much every December. I think the criticism in the article is unfounded and whomever needed/wanted a release should have asked first instead of writing an "angry" blog post.
Author here: I don't need or want another Bun release or an NPM release or anything like that. Like I very clearly say in the article I just got chatting with a peer about the Bun rewrite and I decided to take a look.
I'm consistently skeptical about new tech whether it is NoSQL or Blockchain or Serverless. Some of the things I'm skeptical about fail and some succeed.
The bun Port is a phenomenal engineering achievement that would have taken a team of people over a year or two to deliver previously - look at the TypeScript port for example (and they had llms).
The fact it's taking a while to release is still insanely fast. I'm sure they may be some bumps.
Being skeptical is good and some stuff succeeds and some doesn't - but building tools is still fun and cool, whether it's useful or widely adopted or not :)
Text imports (experimental), native addon ESM support, blob textStream(), ReadableStreamTee, byob for readFile, better event loop monitoring and lots of security fixes as well as more minor bug fixes and improvements. (Last 4-6 weeks)
That's not surprising. Claude code is buggy enough, and releases break things often enough, that I wouldn't expect users to distinguish bugs introduced by switching to rust-based-bun from the normal garden variety bugs.
I can count the number of times an @ file reference doesn't autocomplete on a hundred hands. Or how a rewind won't reset the "is this file Read" marker. Or how a Branch (forking a convo) takes literally 5+secs to run. Or how a slashcommand that's user-only will not work if its in the middle of a prompt. Or how ctrl-r search will match results that don't include any search terms. Or how you can't resume a branch given its session id.
Yes Boris, tell me more about how "coding is solved".
Edit: literally just now, while writing a claude code hook, opus5 gave me this LOC because apparently the changelog and what's in the transcripts _differs_; great docs d00ds.
# Present only for a subagent's calls. Both spellings accepted: the changelog
# names `agent_id`, transcripts use `agentId`, and we need not care which lands.
agent=$(printf '%s' "$input" | jq -r '.agent_id // .agentId // empty')
Claude Code is quite buggy but doesn't generally crash, which is what you would expect if it shipped an immature backend for a month. Maybe the rewrite is full of bugs and they happened to result in the UI glitches or trashed settings files or whatever other application-level bugs that Claude Code has routinely rather than crashes but that would be pretty surprising.
I’ve come to just expect that my CC instance will randomly “blank” and that I have to resize my terminal / use page up/page down to get it to show again.
Supposedly they used a game engine to render their TUI but I’ve never had an FPS game do that.
Ok, so you've just admitted that you have deployed thousands of lines of likely non-human reviewed LLM generated code (Bun's rust rewrite) to millions of client machines (via the Claude Code app auto-update) with significant local client access credentials, in a relatively quick and rushed manner.
What's stopping the LLM from having unscrupulously injected something nefarious into the codebase that you are unaware of?
What's stopping future updates from the LLM from doing the same?
Are you aware of the potential ramifications of deploying thousands of lines of non-human reviewed code to millions of users machines?
Are you happy to be personally responsible for the horrendous outcomes that could occur in these situations?
THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
"I put it in the terms of use, so I am not responsible for what my poorly developer software does" only goes so far.
It could be argued, probably successfully, that this is a case of gross negligence and that Anthropic should be held accountable for harm caused due to their reckless actions.
Especially now that they have been made aware of the possibility, they also cannot claim ignorance of the potential issues.
Certain consumer rights, depending upon the country, also cannot simply be removed or waived due to a organisation's terms of use or implied contract.
I would be stunned if there was a single line of hand written code in the entirety of Claude code lmao. Or if more than 20% of the code had been read by a person at any point. Why would you expect your ai code harness to not dogfood?
CI has always been expensive for Bun including before the acquisition. We build for [macOS, Linux, FreeBSD, Android, Windows] x [ARM64, x64] and then run tests on multiple Linux distros with multiple shards, multiple macOS versions and Windows for each architecture.
We recently started cross-compiling all the builds on Linux arm64 and that made it a little faster (I wrote a CLI tool to download the correct macOS headers for cross-compilation). We also have a daily cron job that asks claude to make the slowest tests faster while adding more assertions.
I think the question for CI costs is still out there. While I do not think it should be in tunes of thousands a day but thats the skepticism presented in the article. True costs are really important to make a good decision in situations like this. One has to consider the fact that lots of people are going to use these numbers to justify the rewrite in future.
It spawns ephemeral EC2/Azure instances, which is a lot cheaper than the GitHub actions runners we used before that.
We shard to a lot of machines for tests and I’d be worried about running out if we used dedicated servers.
BuildKite is fine but I wouldn’t be that surprised if we move off of BuildKite to a custom thing at some point. Months ago, we switched from CMake to a handrolled typescript build system and it made our builds faster and simpler.
codex trivially built me a self-hosted gh actions runner workflow for ephemeral vm's (just a big bash script that manages the vms with qemu). i even sped up the builds with my own custom base image with everything installed in it that i need too.
it works flawlessly.
fanning that out to starting and stopping instances wouldn't be too much of a stretch.
Presumably because they "build for [macOS, Linux, FreeBSD, Android, Windows] x [ARM64, x64]" and self hosting all of that would be time-consuming and expensive.
Probably because they don't want to self-host Windows or MacOS servers when they can pay someone else to do that for them (or Linux ones, I assume that is within their wheelhouse for production but CI is a bit of a different beast to model inference).
It was a commercial slogan for a toilet cleaning agent. An English equivalent would be "mr clean recommends you use mr clean". Today, it is used to point out when someone tells you they themselves delivered good work. Anytime you'd use the meme of Obama giving himself a medal, you could use this phrase.
Actually the existence of the French name is confusing for English speakers as although many would know enough school French to get "canard" == "duck" if prompted, in English this word ends up meaning some sort of fabrication or hoax, apparently from an old French joke where somehow that punchline of that joke was popularized in England so long ago we don't have records. Human culture is weird.
So the first thing an English native might get from "Canard WC" is probably hoax toilet, I expect the duck imagery of the product would bring the "canard" == "duck" meaning across though.
I don't see how it does. Claude Code is an extremely widely used product; the preceding comment offered an objective evaluation target, not a "trust me it's good" argument.
I don't see how Claude code being a widely used product is relevant to the person who orchestrated the Rust rewrite of bun saying that the orchestration of the Rust rewrite of bun went well. Wc eend is a widely used product as well, if that helps.
Ah, for context, which I suspect you may be unaware of, Jarred (the person who said the rewrite to rust went well) is the creator of bun, and the guy behind the rewrite.
I think the chain of reasoning is not hard to follow:
1. Assume the Rust rewrite of Bun went badly
2. Then something must be grievously wrong with a released bun runtime based on that code
3. Claude Code uses the released bun runtime based on that code
4. From #2 and #3, something must be grievously wrong with Claude Code
5. If something were grievously wrong with Claude Code, users would reduce use of Claude Code and use alternative tools
6. From #4 and #5, users are reducing use of Claude Code and using alternative tools
7. Claude Code has wide use and use is growing across all software engineering verticals
8. #6 and #7 contradict
9. From the contradiction, the assumption in #1 is false
The parts that are not explicitly spelled out here are an exercise for the reader. It doesn't really matter if the guy who wrote Bun said this or my uncle said this.
Ok, thanks for explaining! Doesn't really explain how the wc eend expression _doesn't_ apply here (which would be hard to do, because it _does_ apply, since this is someone praising their own work, and it doesn't get more straightforwardly applicable than that), but I do really appreciate the effort.
I always thought it was common sense and just basic critical thinking to take people paid by Anthropic praising products of Anthropic with a grain of salt (and let's be clear, that's what this is), but apparently it's not.
They declared bankruptcy on the original code base so hard that they decided to chuck it all in the trash. I wouldn't hold my breath expecting support for their existing users.
The initial post said something like 11 days and $165k in tokens and it's done. The charts in the OP suggest that Something Interesting happened between July 8th and today, burning way more tokens and human labour. Wouldn't you want to know what?
Either you mean retrospective, or you're being unnecessarily mean for something that's been getting quite a lot update reports, just on a social medium you probably don't use =P
A postmortem is for reflecting on something that went so wrong you had to kill it, and you want to objectively describe what led to that, and what lessons can be learnt from the failure.
A retrospective is just the reflecting part, without implying you believe it is going to fail and you're looking for some schadenfreude ;)
This might be terminology that's used differently in different places.
For instance, in video games it is extremely common to use "postmortem" to mean "post-shipping."
Yes; I think what I described matches that meaning?
Post-death, as in: "it's over," "activity has halted," "no more work being done."
Death does not necessarily have a bad/negative connotation. It's just the end of something.
This is in contrast to the original comment: "something that went so wrong you had to kill it." Death does not imply something went wrong or that it happened forcefully.
I guess "autopsy" generally is talking about "finding the cause of death" so maybe that has a more negative connotation ("something (bad) happened, causing death, let's find it").
> But postmortem is literally “post death” in Latin.
Sure, that's true. It's not really relevant to people who don't speak Latin.
> It’s synonymous with autopsy!
This isn't true. I have no idea how you got here. In English an autopsy is a surgical procedure performed on a corpse for the purpose of identifying the cause of death. In particular, it's a noun. "Post-mortem" in the etymologically literal "after death" sense is an adverb or adjective identifying whatever it is as taking place chronologically after some contextually-specified death.
But, you seem to put some weight on the etymology of a word independently of its current meaning. In that case autopsy "literally means" to witness something personally, as opposed to hearing about it from someone else. (It's "self-eye" in Greek just as post mortem is "after death" in Latin.) Somehow I doubt that's what you had in mind...?
An autopsy is an exam that is performed postmortem. And that’s really how people are using postmortem as a noun for the analysis that they perform after the “death” of a company, so it is a synonym at least approximately, if you don’t see that I guess you don’t have a very good imagination.
Try and find me some sentences out there on the internet in which one of the words could be reasonably substituted for the other one.
They don't mean similar things and they aren't used similarly. A postmortem is, as you note in part, a report or a meeting for the purpose of producing or discussing such a report; an autopsy is a surgical procedure. They're as "synonymous" as the words "cigarette" and "cancer".
Languages change. That word "literally", which you just used, meant one thing for decades (centuries?) ... and now it also (literally) means the exact opposite.
You can be the crabby old guy complaining about how the youth ruined English ... or you can accept that languages aren't immutable, and that the meaning of postmortem is no longer it's literal meaning.
I've been reading people use "postmortem" to describe software retrospectives for like... 15+ years.
Mostly in the opensource, or marketing-blog spaces. The detailed investigation and fix report was a favorite genre of blogpost. And if someone's bragging about the design of an enhancement, or investigating a bug, then the overall product isn't dead. They're (usually) not talking shit to anyone. It's the bug that's dead. Or security indecent that's over. Or a schedule milestone that's "dead and burred".
It's a 1 chili pepper level of spicy to use a slang term that relates to death. I'd never read it as hoping someone fails. I think people let the corpotalk center of the brain overreact to anything that's not couched in euphemism.
> The goal of a postmortem is to draw meaningful conclusions to help you learn from your past successes and failures. Despite its grim-sounding name, a postmortem can be an extremely productive method of improving your development practices.
- How many people have updated Claude/Bun to the latest version.
- How many subscribers care about reporting issues. Most of them are forced to use the tool against their will and have mentally checked out already. Why report issues if your employer values slop code anyway. Just log the hours and keep your head down. Maybe it is not expedient for the AI narrative to report issues!
- How many subscriptions are real vs. bulk distiller accounts.
- If subscriber numbers are inflated.
Judging by the weird Claude Code Github issues page, there are suspiciously few new issues: about 2 to 3 a day only vs. alleged subscriber numbers of 4 million.
I do not trust AI agents to run outside sandbox and use them only in dev container. I always use latest version available when I build container (once or twice a month). In my opinion it is to risky to allow auto update for SW which is released several times a week including weekends and is capable of/willing to do script kiddie pranks :-)
Well, OK, you do that. Many others don't. Not sure what your comment adds except saying "not everybody allows audo-update" which is IMO redundant and obvious.
Jarred, thank you for working on Bun. Many "vibe coded" :D projects start strong and are later abandoned (like potentially Anthropic C), so I understand why people worry about Bun's future. I hope Bun lasts for many years, like GCC. Bun is fast and great to use.
The issue is that the Claude's C Compiler repository does not explicitly identify the project as a proof of concept. It was largely produced by one person directing Claude, and Bun's Rust rewrite also seems to have been driven by roughly one person using Claude Code. That similarity is what worries me. Bun's rewrite could also turn out to be proof of concept.
I guess I'm debating, but it seemed clear enough to me? It proved that AI models and their harnesses are to the point where you can give them some work to do and leave them unattended for a long time, and they'll keep doing productive work for quite a while. This was a novel thing, and quite unclear, at the time the experiment was performed.
Obviously, the word "productive" is doing a lot of work there, but in my understanding the intention was nowhere near "commercially viable" or "practically useful", it was more like "not doing stupid shit like writing comments of the form 'This file contains the implementation implementation implementation implementation implementation implementation implementation implementation implementation implementation implementation implementation implementation ...'".
Maybe somewhere in the vicinity of "either passing more tests or generating more valid tests"?
Grouped and arena allocations work really well in Zig. For awhile, we tried to use this pattern almost everywhere in Bun but it gets really tricky when there’s some GC-managed memory and you want to free things incrementally to reduce RSS. Also, using arenas for arrays that grow wastes memory a lot since it keeps every previous version around (mimalloc arenas are slightly better for this)
Grouped allocations works especially well in parsers & ASTs where the lifetime is very bounded. Since the Rust rewrite, we still use arenas for Bun’s parsers and the bundler but not a ton elsewhere.
I have learned so much reading Andrew’s code and as I said in the original post: Bun would never have happened without Zig.
> The post claims they were fuzzing their Zig code, while during our calls the whole Bun team told us that they were not fuzzing anything. This appears to be an outright fabrication.
If that's what he meant that doesn't speak well for his communication in this post, because "the Bun team told us that they were not fuzzing anything" is a wildly different claim than "the Bun team told us they weren't using a specific tool for fuzzing".
Yes, "their" refers to Bun's code, not the Zig compiler's code. Fuzzili is a fuzzing engine for JavaScript, so integrating it into Bun means that Fuzzili is fuzzing Bun.[0]
From the Bun post[1]
> We fuzz Bun's runtime APIs 24/7 using Fuzzilli, the JavaScript engine fuzzer used by V8 & JavaScriptCore
From Andrew Kelley's post today[2]:
> The post claims they were fuzzing their Zig code, while during our calls the whole Bun team told us that they were not fuzzing anything. This appears to be an outright fabrication.
Sumner says that the Bun team has been fuzzing Bun's Zig code. Kelley says that this is a fabrication. Sumner showed proof that the Bun team has been fuzzing Bun's Zig code.
It looks like Kelley is incorrect and made an unfounded claim. The generous interpretation is that at the time Kelley and Sumner had a more collaborative relationship, Sumner was not fuzzing Bun's Zig code, but I'd expect Kelley to check if anything had changed since then before publicly accusing Sumner of lying in this week's Bun blog post.
AFAIU fuzzing code != fuzzing results. Through skimming it seems that integration tests were using fuzzing, but I would call it fuzzing the code itself.
From "product" perspective there's no difference, but in program-compiler perspective (and e.g. raising bugs about compiler), Fuzilli isn't fuzzing.
Per Wikipedia
> (then...) The program is then monitored for exceptions such as crashes, failing built-in code assertions, or potential memory leaks.
As for myself, I wouldn't use term fuzzing for integration testing such the one used by Fuzilla. I always caught it dynamic testing, scenario testing and in bigger cases property based tests. Fuzzing in my mind is reserved to a low-abstraction calls.
I don't understand what distinction you're trying to draw here. The very specific claim[0] in the Bun blog post that Kelley is calling a fabrication was:
> We fuzz Bun's runtime APIs 24/7 using Fuzzilli, the JavaScript engine fuzzer used by V8 & JavaScriptCore
It does not look to be a fabrication, and is very explicit just about what they meant by fuzzing.
[0] I mean, that sentence doesn't actually match Kelley's paraphrase, but it is literally the only claim in the post related to what fuzzing was done on the Zig-based bun codebase. So it has to be what Kelley was referring to, and his paraphrase is as sloppy as his fact-checking.
For me, using Fuzzilli for testing a Zig code is not fuzzing, it's integration testing. If you're running code externally (e.g. wrapping binary) you cannot guarantee that side effect isn't caused by IO. I consider fuzzing a low level activity with many external variables removed.
Depending on where you are and how you communicate semantics matter more or less. It's very similar to compiler/transpiler. E.g. TypeScript "Compiler" is called compiler but in fact it's transpiler (it emits other high-level language as a result).
My point is that Kelley did not argue that what Bun does isn't really fuzzing. He wrote that the post's claim is a fabrication. But that claim is really specific, and to evaluate whether it is true it doesn't matter what Kelley's unstated definition of fuzzing is.
So an argument about definitions doesn't seem super valuable here.
I'm aware he edited it, but the original version isn't wrong, either. Just like Jarred's
> We fuzz Bun's runtime APIs 24/7 using Fuzzilli, the JavaScript engine fuzzer used by V8 & JavaScriptCore
isn't wrong, even though that was only being done for the last 5-6 months of Zig Bun, and not the previous 5 years when they were accruing all of their tech debt.
If I say, "I run 5 miles every day" and my old neighbor says, "I lived next door to him until 9 months ago, and he definitely doesn't run 5 miles a day," and then I show my GPS logs proving I've been running 5 miles a day for the last 9 months, I am correct and my ex-neighbor is incorrect.
If Sumner had said, "We've been fuzzing our code for years," then Kelley could justifiably say that's incorrect. But Sumner is saying that currently Bun fuzzes their code, which is true, so Kelley appears to be incorrect to claim it is a "fabrication."
That misses the mark here. Every other Kelley's claim has also been unsourced and not cited. I was able to guess that the integration might've simply been something which happened after they'd once not had it.
Given the lack of due diligence here, it seems Kelley's intentions weren't to be objective potrayal of truth but whatever was most damaging.
> For me, using Fuzzilli for testing a Zig code is not fuzzing, it's integration testing. If you're running code externally (e.g. wrapping binary) you cannot guarantee that side effect isn't caused by IO. I consider fuzzing a low level activity with many external variables removed.
I've never heard anyone restrict the definition of "fuzzing" in this way. If I repeatedly generate inputs to a program and then run the program with those inputs, that's fuzzing. It doesn't matter if there's IO or not.
> Depending on where you are and how you communicate semantics matter more or less. It's very similar to compiler/transpiler. E.g. TypeScript "Compiler" is called compiler but in fact it's transpiler (it emits other high-level language as a result).
It's still a compiler. It translates code from one language to another. You can argue whether we need the term "transpiler," but a source-to-source compiler is a compiler.
> You can argue whether we need the term "transpiler," but a source-to-source compiler is a compiler.
That's true today, but compiling was historically was defined as getting source code (human readable) to bytecode (machine runnable without an interpreter).
Some people didn't like that definition, and consequently the waters have been murkied. Just like with eg crypto. Or real time.
How historical? Compilers that have translated from a source language into C and have left the C-to-bytecode translation to another compiler have been around for a long time, as have compilers that translated from a source language to Assembly
I'm not sure what your point is, we all already acknowledged that people have been using the term differently, consequently changing the definition of the word over the years
The conversation amounts to “You should fuzz your code” “we’re already fuzzing the big external dependency, using their own fuzzing setup that they already use upstream”.
It’s not nothing, but clearly not what Andrew meant.
Based on timeline, it seems like both are true. They stopped communicating around the time of the acquisition per OP, which was announced December 3rd, and the PR integrating it is was merged the tail end of November.
Your links show you used a fuzzer, but that doesn't address the other half of Andrew's statement. Is Andrew misreporting/misremembering your conversations?
EDIT: It's really telling that asking a factual clarification question is somehow downvote worthy. I probably shouldn't be surprised, but this epitomizes the reason online discussions devolve in to flame wars (even moreso than real life, though it happens more and more there as well).
The answer could be as simple as we didn't use a fuzzer until recently so both are accurate. I honestly don't know, which is why I'm asking. Yet somehow just asking is triggering to people.
No one's obliged to respond to hearsay by trying to guess what the other party might have misheard to arrive at a false conclusion. It's enough to demonstrate that the conclusion they came to was false, and if the other party would like to defend themselves they're welcome to explain why they came to that misunderstanding.
You want to see him address being "a stinky manager", having "beginner energy", choosing to take VC as opposed to "a solid living via crowdfunding", or "already writing slop well before he had access to LLMs"?
Those are all opinions where arguing about them isn't going to be productive.
Countering the accusation of "an outright fabrication" on the other hand is worthwhile because it's a claim that can be countered.
If somebody called me a liar for something that demonstrably wasn't a lie I wouldn't let that stand, either.
> Countering the accusation of "an outright fabrication" on the other hand is worthwhile because it's a claim that can be countered.
Hmm, fuzzing integration was merged 8 months ago. First found bug mentioned 3 months ago. Bun is 4 years old. I think both arguments can be true at the same time based on this evidence. It is entirely possible that for more than 3 years team has said that no fuzzing was done, and the first fuzzing was done just 3 months ago, and this information did not travel.
I agree that ""an outright fabrication" is a bit too much.
But also the claims about the fuzzing in the original blog post are kinda too misleading. Fuzzing harness is basically just coverage-guided random bytes towards Bun's JS APIs and it will not really catch anything in depth from the code. Just the most obvious from the surface. And 24/7 fuzzing is introduced likely around the same time when Rust rewrite seemed to be main focus, as then the first issues were created. Current fuzzing does not give much trust about the code stability, but indeed the fuzzing has been started and likely improves in the future if someone puts some work for the harness.
> And 24/7 fuzzing is introduced likely around the same time when Rust rewrite seemed to be main focus,
This is false, it was done many months before (Nov 20 2025). There is some irony in your comment being in reply to a thread about verifying claims before posting them...
> This is false, it was done many months before (Nov 20 2025). There is some irony in your comment being in reply to a thread about verifying claims before posting them...
There is no evidence for this. It is integration PR to support fuzzing with one tool. The public repository does not have CI pipeline for it. Fuzzing is done privately somewhere, and the first linked issues are around April.
Those statements are all consistent with the publicly observable facts like this whole thing going from "I'm experimenting with this" to "this is merged into main now, yolo" within a week or so. The complaints from Jarred on the amount of bugs they've been having is too.
Bun donates $60,000 per year [month] and Jarred acts with graciousness and soft tones when talking to outside parties. Why do you think it's Jarred's obligation to continue after being insulted for professional dishonesty?
Isn't it Andrew's obligation to show that he was worth that much kindness to begin with?
Maybe I'm overly cynical, but I don't know that directing some funding towards the open-source project that is the foundation of your whole tech stack is really "kindness" per se.
When you are vending a devtool to other open-source developers, and making a lot of hay about the specific technology choice, it's basically marketing spend. It's also often a way of buying favour (attention to issues, PRs, etc) from the project maintainers...
Without having any opinion on whether or not the Bun team was meaningfully fuzzing their codebase... Andrew's claim was not about whether or not they were, it was noting that the story was different between what they claimed in conversation and what they stated in this article.
All of those commits except the initial integration are from after the acquisition. How do we know this wasn’t done without anyone on the call’s knowledge?
Since there are multiple ways to interpret Andrew's original comment, multiple ways to interpret what his newer edit of it implies, multiple logical reasons each way could have come about, and likely multiple opinions on what the expectations from each side should be... I'm finding myself getting stuck in a loop of trying to understand how/why these considerations are important.
Could you help further explain which interpretations & ways you feel this info is relevant to?
This PR is an implementation of the design from https://webkit.org/blog/7846/concurrent-javascript-it-can-wo.... I think it would be really cool if JavaScript had true shared object multi-threading without compromises (SharedArrayBuffer, postMessage are not that). If we had both threads and structs, it’s likely the TypeScript compiler would never have needed to be rewritten in Go.
The title should be changed to clarify that it’s a PR to Bun’s JavaScriptCore fork and not the upstream WebKit.
This PR is scarier to merge than Bun’s Rust rewrite PR. There are a good number of benchmarks/stress tests, unit tests, and also TSAN runs and security scanner runs, but this is a more complex change than the Rust rewrite (yes, really). I’m also worried about syncing with upstream - today the “fork” is mostly a bunch of patches, but with this PR, changes to the JIT need to be reviewed for behavior when multiple threads are in use. Our best bet for this to move forward is figuring out a way for some constrained version to be upstreamed into WebKit proper, if that makes sense and if they’re interested.
I don't find this problem any different than with human writers. Agents are verbose, sure, but I mostly find them providing far more useful information (in far less time) than your average (P80, really) SWE.
Yeah, I know. Not a single comment in reply to any questions of substance in at least a month. I suspect he knows that his choices are indefensible, that's why he doesn't bother to defend them. Can't wait to read that blog post that's totally coming any day, though.
This wouldn’t fly in college and it won’t fly for me. Claiming authorship means that you did the work. If you just prompt claude, then you authored the prompt. Not whatever claude slops out.
Jarred should be putting the prompts in github, since that’s the only original work involved.
And to be clear, I don’t think that being responsible for what is produced and claiming authorship are the same thing. I do agree that Devs are responsible for what they ship.
I've seen the Bun Zig->Rust MR a few weeks ago when it was current. Now I'm seeing this, and I have to ask, since you're here:
Is there no way to make this changeset smaller?
At work, I've usually written large patches. I used to be worse at it. I was mentored out of it, and while I still like my patches to be complete, I balance that with the available bandwidth of the team and what the team can reasonably actually process.
For perspective, my "large patches" were PRs on the order of 10-12kLOC for relatively big features. I consider those to be on the upper end of what is reasonably reviewable by a small, non-dedicated team, and towards the upper limit of the kind of PR where I can speak for nearly every line of code, what it does, and why it's there.
On the other hand, now, LLMs are part of the equation, and they can (and often do) write code in insane volumes. They arguably tend towards extreme verbosity, without even talking about docs/markdown files. While LLMs are part of the workflow, my company, and those my friends work at, have all instituted policies of the developer attaching their name to the code ultimately being responsible for the output (which IMO is a lazy strategy, but I can't think of a much better one under the circumstances).
I cannot, personally, fathom how you can stand behind a single changeset spanning 2000 files and a quarter-million lines of diff. Do you consider this sustainable?
At this point the code bases are very quickly getting away from us in the open source community and even in proprietary code bases, and these are important code bases. Often very complex, often legacy. Who ultimately still owns these? Who's really going to be accountable if things go wrong?
Ideally, you could find pieces of that 10k that would also work as a standalone improvement to the app. I understand that team cultures can vary, but doing small features or refactors in service of a larger goal is nice, because while the reviewer knows you must be up to something, they can still approve your preparatory PRs as face value beneficial.
I have some disagreement here. There are factors other than just code review to worry about when implementing changes, quite notably quality assurance.
It does not one any favours if your 10k LoC gets split in 5 changes that aren’t supposed to regress anything (but need to be validated not to) then 1 tiny on that brings things together.
Some features will be confusing for end users if you drip feed them. We had a whole host of changes recently overhauling our moderation system to be able to track and audit compliance with DSA and the key factor is ensuring the system makes sense to our users and that they can enable, have documentation and on-boarding materials for the changes in functionality and that it’s all QA tested.
In this case we did still review smaller chunks of code, we accumulated them into 1 large merge request at the end and merged it after QA.
That’s a good point; it depends on the extent to which you can make either invisible changes or you can roll out improvements that don’t require coordinated communication.
It's easier to justify in a fast-moving greenfield code base with a verbose language... but I won't defend it. I've gotten better and I'm still getting better at breaking these things up.
I brought the 10-12kLOC PR up as an example supporting my point of view. I don't encourage the behaviour. Most of my PRs these days fall under the 1500LoC mark, tops -- maybe a bit more if it's a tricky component that needs a ton of tests.
10k in 2hrs is 1.5 lines of code per second for 2 hours straight without spending any time to make comments, think about what the code is doing, etc.
In pre-ai era that is just skimming and trusting the person who wrote it or the code changes are largely auto-generated or there exists an exceedingly simple test suite that is incredibly verbose.
Post-ai you are ruining your code base, I probably have to spend 3-5x longer reviewing ai generated code, the code they write tends to be too verbose, mediocre, filled with subtle bugs, adds unnecessary comments, etc. If someone gives me 10k loc pr it's a sure thing they've just let the ai run loose and I'd just tell them what they need to change in general terms instead of wasting days of my time reviewing junk.
The point is that there is a 0% chance you can meaningfully understand the impact and consequences of what a 10k line change is doing in a couple hour review, so there’s no way for you to know if it’s “bad” in any sense other than it’s so awful that it’s obvious on a skim (which is what 10k lines in a couple hours is).
> I was mentored out of it, and while I still like my patches to be complete, I balance that with the available bandwidth of the team and what the team can reasonably actually process.
I struggle with the same issue. In my experience you can't reduce the total number of lines. If the feature took 10k or 15k loc to deliver it, you aren't going to be able to reduce that meaningfully.
You can usually break it into stacked series of commits. New code can be split up into stand alone modules, which all compile and pass their unit tests. They could even be shipped, although they wouldn't be used because the changes to the UI are always the last piece of the puzzle. If you are refactoring, you can usually find a way to split the refactor into smaller steps, each building on the previous one. That is almost certainly possible in this case.
The issue with both approaches is while you can review each step independently, what you miss by looking at just that commit is the motivation for doing it. You can only get that from the big picture, and to get the big picture you need the entire 10k or 15k loc available.
That means you have to push the entire series of commits. If you want to make it plain they are individually reviewable you push them as stacked PRs. Either way, it's a 15k loc push.
I don't see a way out of that for the same reason neither bottom up nor top down design works on their own. You have two edges - the upper (often the UI), and the lower (the OS, standard library - things you have to use to get anything done). You work from both edges simultaneously, each working towards the other, hopefully so that when they meet in the middle and the two fit together nicely. The point is, you have to review like that too. You can't just look at how neatly the blocks are stacked on the foundations, you have to evaluate if they are taking the best route to the destination. The review can fail in both ways - the UI can be beautiful but it stands on a mess, or the code could have built up a beautiful series of abstractions that bubble through to the top level and ultimately confuse the end user. So you have to review the code from both perspectives, and to do that you ultimately need to get your head around all 15k loc.
This means a reviewer demanding they be spoon fed a few thousand lines of code at a time is being as unreasonable as the person delivering 15k loc in one commit. They are demanding a simple solution to their problem, and it is wrong. They should be demanding all 15k loc be delivered in the form the author intends to ship, but split into digestible commits that clearly explain the path and reasoning taken between the top and bottom edges, so both the top and bottom level designs are plainly visible.
What happens when I do that is I get into fights over forced pushes. Everyone hates them, and for good reason. They asked for a simple change in their review, and what they want to see is a small commit reflecting their request. Hiding that by doing a rebase is met with howls of pain: "no forced pushes!". So you insert your commit reflecting just the change they asked for into the stack of commits the large feature necessitated. Doing that rebases every commit that depends on it, of course. You push the result and are treated with a chorus of "NO FORCED PUSHES".
Forbidding all forced pushes makes about as much sense as forbidding a 15k loc change, even through its well-structured into commits. It makes me wonder if unis bother to teach modern software engineering practices.
I will say that a good solution to this starts before a line of code is written, and does require a PM or scrum master with a deft touch (ideally one who's been involved with the engineering side of a project for a while).
Working with them, the scope can be brought down, sometimes, though this obviously depends a lot on externalities like who the customer is and how much flex they're comfortable adding to the timeline, since that often becomes a factor.
Part of becoming a more mature developer involves being able to navigate these situations with PMs better, and dealing with the frustrations that can often bring (in my experience). I'm still working my way up this side of the job. Historically this stuff has been managed up really thoroughly by my immediate manager (a factor of both the manager's working style, the work being done, and the broader company structure), but my current company structure means that I have to get better at this stuff a lot more actively.
>mostly share read-heavy graphs and coordinate through a few hot objects, which is what Lock/Atomics are for.
Then it is a clear overkill to me. I’d rather built an in-memory DB on top of shared array buffer. Would work almost as good as an object graph but does not require a full system overhaul.
To be clear, that was my opinion as a delegate, not the consensus of the committee, as that hasn't been formally discussed. That said, I believe other delegates share my opinion that shared memory concurrency needs to come with more constraints than what is suggested here, and while some delegates might be fine with it, I doubt any unconstrained shared memory concurrency would ever reach consensus if brought forward as a proposal. Currenntly the language spec is actually unclear on whether this is allowed for a compliant engine, and something we might want to update.
I have respect for past you, whose accomplishments are incredible. Current you I think of like a circus performer. I trust you to get ooohs and aaahs from the AI-race spectators who don't know any better. I don't want someone who is primarily in showbiz within 10 fucking miles of my infrastructure though
Is it possible to merge it, but keep it disabled by default? This could allow users to play with it on Bun while maintaining the expected behavior of JSC.
Dynamic workflows, in my experience, make Claude more effective at complex long-running tasks. They help precisely with getting Claude to do the task correctly.
It feels more like a bespoke build system for the specific task/project than prompting a freeform chat.
As long as agents are fuzzy (which they will continue to be with the Transformers architecture), the need to validate will continue to exist. I cannot imagine merging code without at least 1 human review.
`PathString` worked the exact same way in our Zig code, with less visibility from the compiler & type system. And yes, it will be refactored heavily (or deleted overall) in the next week or so.