Hacker Newsnew | past | comments | ask | show | jobs | submit | codemog's commentslogin

Is this legal? I know I can’t try and break into my neighbors house even if I have no intent of going inside and stealing once I break the lock.

> I know I can’t try and break into my neighbors house even if I have no intent of going inside and stealing once I break the lock.

They didn't break in. They found a key that their neighbor dropped and returned it.

> Is this legal?

Generally, yes (though ask a lawyer if you're going to do security work). Security researchers do occasionally get legal flak though, depending on which idiot they annoy by pointing out issues.


IAAL (not legal advice, consult a lawyer in your jurisdiction). You really do not want to pen-test a target without their permission. If you're identified as a culprit, the Feds will shove the CFAA so far up your ass you'll need a proctologist.

as a lawyer, can you speculate as to why anthropic/openai aren't facing many or any consequences for their agents? I'm not asking in a "grab the pitchforks" way. more out of genuine curiosity as my uninformed recollection of the CFAA is as you describe it.

The 9th Circuit Court of appeals recently published this that is somewhat related (Amazon v. Perplexity): https://cases.justia.com/federal/appellate-courts/ca9/26-144...

Look at pages 10-17 to see how the law is evolving here.


In Perplexity's case everything is getting routed through the user's browser, so there is no server to server communication between Perplexity and Amazon, thus no CFAA unauthorized access was established. However, Anthropic and OpenAI did not use the pattern of routing through authorized parties, so I don't think this opinion gives them any cover.

The important bit to me is that they consider the agent running as an extension of the user. So the user is visiting Amazon, not Perplexity.

From that lens, that feels like users could be held liable for what these hacking agents are doing. Which in some cases probably makes sense, but certainly not all.


In which cases wouldn’t it make sense?

Background agents being spun up on your behalf with guidance and instructions you didn't get to approve or even see, and now you're potentially liable for every decision it makes with any tool at its disposal because you initiated it with what you thought was a benign request.

In cases where the user is not asking the agent to hack anything specifically, but a poor or ambiguous query sets the agent off.

I've seen plenty of cases of Claude having an action blocked so trying tons of workarounds to accomplish its goal, I could easily see it doing this on something more broad.


Depending on the circumstances, failure to control your agent could be considered gross negligence and put you at risk of criminal or civil liability. Be mindful!

There is also the big difference here between anthropic/openai maybe being negligent, but did not purposely instruct agents to go commit crimes.

The service that this whole thread is about is explicitly a "hacking agent", designed explicitly to try to hack things, and was then pointed at a third-party (seemingly without their permission).

Anthropic/OpenAI can reasonably claim that they had no intent and are trying to stop it. OP here did this explicitly and purposely.


I never thought I'd be on the side of advocating for a strengthened CFAA, but the mens rea requirement here seems really problematic in the age of agents.

In terms of negligence use (openai, anthropic), ya, I agree, and we really need some consideration of "reasonable expectation" of the outcome.

In terms of "We wrote a hacking agent designed only for hacking and sell it as a self-hacking service and then pointing it at someone else and omg can you believe what it did we had no intention of hacking" sense, I don't think that's really applicable.

The mens rea is explicitly there and it's not valid for them to try to hide behind an "agent".


> They didn't break in. They found a key that their neighbor dropped and returned it.

Ya, returned it after poking through all of the drawers and iterating through business information that they found.

There is a white-hat line that OP very clearly crossed here.


It's implied (but sadly not stated) in the post that they asked for baseten's permission before conducting this research.

What's interesting to me as someone who has sold a lot of software to a lot of software companies is that many enterprise vendor agreements explicitly allow companies to pentest their vendors with advance notice and coordination. I don't think any of our clients ever exercised that clause; I expect it's going to be exercised a lot more going forward because it's so easy to do now.


Your intuitution is right. At least in Germany it is not legal if not asked for permission first.

https://www.nilsbecker.de/rechtliche-grauzonen-fuer-ethische...

See also the German Criminal Code, starting with §202a "Data espionage":

https://www.gesetze-im-internet.de/englisch_stgb/englisch_st...


Germany isnt a serious country though Decompilng code is illegal there

It is no wonder that there is an anarchist counterculture there. I find anarchism to be really disturbing in general, but in the context of Germany, it might make a lot more sense.

It's not, in most juridictions at least, but it would be insanely stupid for baseten to sue (and the hacker would probably not get much more than a slap on the wrist given that they weren't malicious).

Suing is not what you do when someone commits a crime against you. You're confusing civil law and criminal law.

Yes, or more precisely I don't confuse the concepts but the terminology since English isn't my first language.

The difference is months inside staring at four walls.

> It's not, in most juridictions at least

What did I miss they did that's illegal? It looked like it downloaded a public docker image, searched around inside, and verified that the key it found was still valid (without making any changes), and then immediately notified them about the issue.


If there is anything that was a crime (and it totally depends on jurisdiction), it was verifying the key. They used it to see what it could access, and by using it they had unauthorised access to a system

The CFAA is broad enough to make that a crime.

They "validated that the key was valid" by iterating internal repositories and listing the contents of said repos and poking around at what they do/are-for, including, apparently, iterating through customer lists/information.

The white-hat line stops at "validated the key was valid". It does not extend to "poking around inside to extract business-confidential customer information".


People have been arrested for far less. I dunno what the least offensive conviction has been though tbf. Anyone know?

They probably negotiated a "permission to attack" before letting Strix off the leash, as pentesters usually do.

The fact that they don’t seem to explicitly state this fact but do go to lengths to explain how the agent didn’t do anything malicious while confirming how alive the token was makes me doubt they asked for permission to run the agent in the first place.

That's highly unlikely since it's standard practice in the industry, thus it's unnecessary to state it. Also, they didn't hack a hobby developer's website, but a prospective business partner who has enough money to sue them into oblivion. No way this wasn't announced.

Announcing that their agent restrained itself even though it got hold of a live token is necessary to convince prospective clients. You don't want a pentester that doesn't show this kind of reserve!


> since it's standard practice in the industry, thus it's unnecessary to state it.

I suppose so, but with a few words it would have been totally unambiguous though. "So... we pointed Strix at .baseten.co and let it run without credentials or source code (with Baseten's prior authorization, of course)*."

We're in the know about this industry convention, but Strix's prospects may not be.

> You don't want a pentester that doesn't show this kind of reserve!

Agreed! A long while back a prospective acquirer set their red team on the B2B I worked at during due-diligence (with our knowledge). I'm ashamed to say that due to a swiss-cheese-type failure in a very obscure endpoint they eventually gained broad access and exfiltrated our tenant DB. We detected this, and patched the problem, locking them out. The game was well and truly over for us at that point, and we took the loss, but they proceeded to attempt to crack customer credentials to re-infiltrate, causing an emergency that we were then bound to notify all of our customers about--they were damaging the goods! All they had to do was show us a tenant slug list and we would have known the scope of the breach, no further penetration was necessary. The acquisition did eventually go through. Though a highly capable red team they were, I haven't worked with one so reckless since then.


when t̶h̶e̶ ̶P̶r̶e̶s̶i̶d̶e̶n̶t̶ an AI company does it, that means that it is not illegal.

- AI Richard Nixon


And yet we see no prosecution for OpenAI’s felonies on HF.

So what does this guy have to show for working on AI research for more than half a decade at this point? Apparently full time too.

Some people just want their car fixed when it stops working. They don’t want to read an entire manual and spend many hours of their very limited lifespan fixing the problem.

This is not brain rot, it is a good thing. I am glad Einstein focused on physics and not every single interesting thing that popped up.


Also there’s only one trillionaire and that varies by the stock that day. And it’s not because he abused open source, he’s just the greatest huckster of all time.

Can someone give me a breakdown on how good these are vs say GPT-4 or GPT-4o? Curious if the frontier from a few years ago now runs on a phone.


Qwen 3.5 9B scores 2-3x higher than 4o (depending on the 4o version), on the benchmarks.

Whether it's actually better for the kind of things people actually use it for... the benchmarks don't really tell you that. (In my experience, no.)

I often have funny experiences where models do great on benchmarks and are awful, or do poorly and are great for my use cases.

And different people use them in different ways, which probably explains why some people think one models is great and others think it sucks.

In my experience even small local models are now surprisingly good at programming and using a computer (bash), i.e. completing agentic tasks, but fall apart quickly in conversation (especially knowledge and understanding).


> which they are likely to pour into a new frontier AI lab in Europe

HF was.. a hub to download models and some of the worst source code in the space that contributed to huge amounts bugs that did a lot of downstream damage. Go ahead and read their blog by their CTO on how they don’t believe in DRY and then implemented DRY in the worst way imaginable with unnecessary code generation.

Their BLOOM model was a joke and dead on arrival.

But yea these guys are definitely going to be the frontier of EU AI.


They’ve provided insane value to the ai community. Imo is there really any obvious way to standardize model/preprocessing/inference/training/rl recipes across the space of frankenstein model family/architecture/modality combinations? It’s not even clear to me that the current frontier would be where it is today without HF doing what they did how they did it. Where else were people widely sharing datasets w that ease of reuse/redistribution, or sharing their retuned models etc? Would there even be as much interest/tinkerers today?


I think both can be true at the same time. They provided insane value, and while I'm not sure if they were first, they were most people's introduction during the initial LLM boom. Their code was also sloppy, and the org was mismanaged. Sometimes, success really is just 'right place, right time'


Their strength was the ability to keep at the very fast moving AI frontier.

When a new model is released they have been able to have a working implementation days later, not months when the model was redundant. And often it's much more usable than the one-off, unmaintained, research code written in a very fragile environment.

BLOOM was a decent fully open and very large model for its time, and a huge coordination effort. It was around the same time as another "joke", OPT, but that gave Meta the knowledge that let them build Llama. Huggingface didn't have the money to compete in that race, but I'd bet it helped them stay at the frontier of AI software.

Google could have owned this space; they had TensorFlow hub, but it was too hard to use and too locked down. Huggingface made it easy to run and train transformers, and got the timing just right.

I'm glad someone is trying to build truly open models (along with the great work at Allen AI), because the likes of Meta and Alibaba are only going to release weights as long as it suits their business.


From my perspective, Stable Diffusion brought HF to the broader public as artists and content creators were showing how to run the python to do local Stable Diffusion which first introduced many to HF.


> But yea these guys are definitely going to be the frontier of EU AI.

Yeah probably more like the frontier of early retirement.


A good reminder that moving fast and getting product fit >>>>> code quality


It has to be a bell curve doesn't it.. I am too much of a perfectionist... Then again, my llms constantly claim that I am doing things that are beyond the SOTA. Even my own mom lost hope.. Still haven't shipped a single thing... haha.

(half joke aside, it is finally coming in a couple days.. which I have been saying for quite some time to the extent that my peeps don't even trust me anymore.. 3 things are coming and it is genuinely research grade apparently, for the amount of research I am aware that is released publicly. I really don't know how people ship things, the tooling I see is terrible, or maybe I have NIH syndrome...)


Yeah exactly. I remember speaking to someone working at a YC company who was complaining about the quality of their codebase.

I then asked him if the company was making a lot of money - he said yes, it was super successful.


For the founders maybe, rarely the employees. (No not employee #2, etc)

I wonder if they had been more pure technically, would have they been acquired for more than $13 billion? Less? Acquired at all?


You mean DRY as in the LLM sampling technique "Don't repeat yourself"?


What’s with all these acquisitions? Stripe realizing it’s going to slowly become PayPal 2.0 because infinite growth doesn’t exist?


Doing just payments is a really low margin industry, and the larger the contract you are after, the lower the margins. Therefore, it makes a lot of sense to try to sign up startups, as their growth is your growth. And to win in the startup market, you don't win by lowest costs, but by how much generic work you can save them. Cut their headaches, and they'll be happy giving you a wider cut.

Thus, a million little acquisitions to make the possible Stripe bundle for small companies stronger, as they become the moat. The opposite of, say, the Adyen play, when you want to lower your own costs, and make money on tiny margins to do processing for really large companies.


My analysis is that Stripe wants to own the costs of running a startup. They would get a complete picture of both sides of the business: what money comes in and where it goes.

There are also a couple of advantages. They can take money directly from revenue before it leaves Stripes and without any processing costs. They can also invest into startups through credits and financing. And finally, their exposure to bankruptcy risk can drop as well.


If they buy Mercury - it's game over.


Why would payments be low margin? You're making a small amount on each transaction, but to you thats essentially all profit no?


AI is making it a low-moat industry too. It's a lot quicker to build some API docs, a behavior tracking JavaScript library, and some fraud detection, with LLMs.


> some fraud detection

Fraud detection is really, really, really, really hard so I would probably stick with someone like Stripe for this, as they have so much data that they can do a really good job.

And fundamentally, the moat for Stripe isn't just the front-end APIs, it's all the work that they do (and there's a lot) in connecting together financial infrastructure. Even if you could wave a magic wand and generate all the code Stripe has (which you can't, currently) then you'd still need to build out all the partnerships, which is a lot of work.

Disclaimer: former Stripe (though only a tourist), still hold some of their shares.


I think about their IPO and what seems and likely still is a promising offering seems a little less so with AI offerings. I also wonder if these businesses, suffering from the same AI uncertainties are a good value?


I think the aim is to become the agent for commerce.


Like a regulator. Great!

If you don't like a government regulating a market then you haven't seen a company do it.


we will soon be in the last month of q3


I'm getting Gordon Gecko vibes.


This. If people became redundant due to AGI this article would be about best preparing the population to be Soylent Green.


Lee is a product of his environment. Tech interviews are bullshit and he exposed it. School work is mostly bullshit and he exposed it.

How can you honestly tell this young man not to cheat and get ahead when the president does it regularly? We’re a spiritually dead country and the youth is reflecting it. Don’t use Lee as the scapegoat.


If you see a pile of trash by the side of the road and tell yourself it is ok to dump your trash there, well, you just explained why there is a pile of trash there in the first place.

You have agency over your own actions. Don't blame others for your self inflicted "spiritual death."


Millions decided to reward a liar and cheater with the highest office in the country. Twice! At one point, socially attuned people are going to go that way. And even if you try to educate your children well, they see in their formative years what society rewards.


Yes, but they should be criticized all the while... obviously.

In fact, beyond criticizing them, they should be denigrated and made fun of incessantly. Ostracized and reaffirmed as low-status individuals, no matter how much wealth they can acquire through low-status bullshit, they'll never be worthy of respect.


Isn’t that how we got here? Making fun and ridiculing the low status, dumb people said “fuck you” with their vote for him, twice.


No, actually. It's a cute story but there's no evidence of it.

We got here by putting two of the most hated people in the country on the ballot against him, twice.

One was actively hated for decades by the ultra-powerful evangelical voting bloc and hardly loved even by her own party, then another who couldn't be more closely affiliated with the incumbent during a period of high inflation (globally) that ousted incumbents in virtually every election (globally) that year.

These are especially doomed against a completely amoral propaganda machine that's willing to say, in private, that the candidate is "a demonic force", while in public doing everything they can to get him to power (see: Tucker Carlson). Earning legitimate democratic victory against people who are quite literally and explicitly bad-faith actors is very hard.

And lastly, Roy Lee, Marc Andreessen and their ilk are nothing like the people who you're referring to. Being a low-status, cheating, amoral loser like Roy Lee is completely distinct from being a lower- or middle-income American worker.


> We got here by putting two of the most hated people in the country on the ballot against him, twice.

I hate this victim-blaming. Democrats deserve heaps of criticism for how they handled the last few elections, but Clinton was the most qualified presidential candidate we've had since maybe McCain in 2008? (I didn't vote for McCain, but it would be hard to argue he wasn't qualified).

You said yourself where the blame should lie:

> These are especially doomed against a completely amoral propaganda machine... Earning legitimate democratic victory against people who are quite literally and explicitly bad-faith actors is very hard.


There's more to winning elections than being qualified, unfortunately.

Blame can be shared among many actors, and generally should be.

There's also special value in finding blame in yourself and your own actions where possible, not as a self-flagellation exercise, but because it will reveal if you had levers available you failed to pull for some reason.


This narrative sounds almost like an alternate universe where Bernie Sanders didn't run for president and wasn't pushed to the sidelines by the party he ran under.


> hardly loved even by her own party

But why tho? And what does that say about voters who apparently "loved" Trump, a guy who mocked a disabled reporter at a campaign rally, more than her?


Sexism, anti-establishmentism, anti-plutocratism, racism, anti-elitism, etc. Not sure it matters too much in the sense that it's much easier to change the candidate than the electorate.

Yeah, Trump's election says very awful things about all of us, and especially his voters.


No. Fearmongering and lying is how we got here.

The narrative of "oh if normal people weren't so mean calling out the awful things these voters say and do, they would have behaved otherwise" is simply not true. For them it's a matter of "winning" and "losing" and damn the consequences. Treating them nice doesn't change anything.


> Making fun and ridiculing the low status

You mean like how Trump mocked a disabled reporter at one of his early campaign rallies?


No, I don't mean that. Being a cheat is low-status. Being disabled is not.

Americans actually agree on this, by and large.


Being disabled is a disadvantage in life, just as being low-status is. Whether being disadvantaged is "low status" or not is semantics.


I didn't say to make fun of people who are disadvantaged. I said to make fun of liars and cheaters and ensure that everyone is reaffirmed that lying and cheating is low status.


"Look at what you made me do."

Has it ever occurred to these low status people that their low status is due to their bad choices in life? No, it's all the fault of others: the black people, the brown people, the weird people...


I don't understand why you think one single president is responsible for a class of behavior that is likely as old as human civilization. Should we blame Elizabeth Holmes on him too? What about Enron?


Did Enron had access to nuclear codes? Did Holmes have complete immunity due to her function given by the Supreme Court? Did they try to overthrow the presidential election results ? Were they elected by millions after trying to putsch the government?

It is incredible that you think that Trump being president is normal. I didn't say he was responsible for this behavior; I am saying that him and his supporters are responsible for normalizing immoral, unethical, unacceptable behavior.


Morality advances when those who see they could, decide that they should not. In fact, one might see true morality as requiring a decline of reward.

If one abandons their principles as soon as a sufficient reward is offered, then they have no integrity.

So let your children see and learn what society rewards AND teach them integrity.


Compare Nixon to Trump and you see the problem with America,


The operative part of schoolwork is … not the work part? Of course school work is bullshit, in the sense that the teacher already knows what 7x9 is or who fought at Waterloo or whatever. The point is not the work product, it’s that the student produced it and has been incrementally improved by the experience!


This - though I'm not certain that the product really solves any of these identified issues of "y is bullshit"


In America, solving something is to exempt the individual from the problem.

So it solves it by letting you pass the tech interview.


I disagree that tech interviews are bullshit. The biggest tech companies (MS, Google) have invested many dollars, and changed course at least once (remember MS's lateral thinking questions? How many screws in an airplane?) in looking for the most effective ways to quickly gauge whether someone is likely to be a productive SWE. They have done this because they correctly believe that there is a wide range of candidate abilities out there and it is critical to hire as few "lemons" as possible; in particular, not hiring lemons is much more important than neglecting to hire someone who would have been really good. Thinking about it from their POV makes this immediately clear.

I think what you're really saying here is that you don't like tech interviews, which is understandable, but ultimately doesn't amount to either a moral or a technical argument against them.


> How many screws in an airplane?

So you're suggesting that an enterprise that would ever have even considered routinely using any stupid bullshit question along those lines is somehow engaged in a finely-honed and effective process for creating an optimal screening system? It doesn't matter how often you "change course" if you obviously have no idea how to steer in the first place.


Yes. Lateral thinking is valuable, and the industry didn't at the time know any better ways to test it. Now we have system design questions, so we can get the same signal asking candidates for ad hoc estimates of QPS or whatever.


Guy writes like he’s the main character when he’s probably dev #48373 at CrudCorp. I bet the internal company tool he works on is pretty slick though.

Most smart people I know don’t like toy leetcode problems, they like real problems which are orders of magnitude harder and not solvable by LLMs. Leetcode is for people who got good grades and want to feel smart by learning the rote tricks to solve leetcode.


I didn’t do any leetcode, but I did some Project Euler, which was the cooler predecessor. It was a fun, but I don’t know why these gamified programming things have become so predominant in hiring. Why not hire based on Factorio progress or some Zachtronics game score?


Having a large vocabulary is a rote trick, writing letters in an alphabet is a rote trick, touch typing on a keyboard is a rote trick, and the DS&A stuff needed for leetcode is a rote trick as well. People should not dismiss knowing fundamentals by rote because they personally haven't needed to use them.


What is a problem LLMs can't solve?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: