Hacker Newsnew | past | comments | ask | show | jobs | submit | bfbf's commentslogin

But people who are bad at maths are unlikely to be writing about maths. A crude example might be if you search for “2+2=” in the training data, you’re much more likely to find “4” as the next character.

Obviously llms are far more complex than this, but I think this proves the point. The fact you had to add the “recent” qualifier there highlights that llms in general were bad and had to be provided with corrective targeted training data to improve. (And they still can’t count the R’s in strawberry!)


> they still can’t count the R’s in strawberry!

My resident Qwythos-9B counts 3 "r"s in "strawberry". So even locally hosted models are catching up.


> And they still can’t count the R’s in strawberry!

Really? I do not have the time to survey the modern LLMs to see if your assertion is correct, but if it is then I'm surprised; I would have thought that that one would have shown up so often in their training data that they would be able to answer that question, even if they would then be unable to (for example) count the R's in raspberry, or in some other word where "count the R's in _____" was not widely found in recent online discussion.


I mean a very quick check in ChatGPT got it wrong today, yeah! I’m not sure what model I was using, but it proves the point!

But the OP already addressed this-he can just ask ChatGPT directly how to fix the brakes. He’s gone looking for a human expert because (for whatever reason) he doesn’t want that.

For me the issue is very clear: I happily use AI to answer tons of questions each day, but if I’m reading your website/blog/article/Jira ticket, I expect real human input.

It should matter whether the author “sweated over it for hours”, because that means they actually considered the best way to communicate something. If they didn’t do that, then they clearly don’t care enough about about the quality of their output and I shouldn’t spend any time at all on it. In fact it’s worse than that because if the author doesn’t respect the reader enough to craft their output, then I as a reader have no respect for their work, and I resent that they’ve taken my attention for the time it takes me to realize it’s ai generated. If you’re publishing Ai content, then I can only presume you have an ulterior motive than plain communication- in the case of OP I’d assume this is for ad revenue or clicks.

There are many types of writing, and I struggle to think of many where a statistical model could feasibly produce an acceptable output for the given intention.

Eg. A poet chooses their words extremely carefully; a good instruction manual is written by the designers of the product (not just guessing at a common or plausible method); a news article has an angle/story beyond just the facts.


If you or OP suspect (suspect) that a blogger's taking shortcuts, and that's a problem for you or OP, then you or OP can just skip that blog and find another that appears (appears) to meet some criteria for labor input. To the extent that this is all unfalsifiable, it just should not matter.

I'll go further. Even if it were falsifiable, why should it matter? An example: I subscribe to a reputable US publication and I enjoy its journalism. If it turned out that one of its writers had (somehow) used AI to generate their article (which I enjoyed) in a click, here's what I would say to them: Hats off to you! How did you do it?!

I don't care if the article took them two minutes or if they used AI any more than I care if they used a spellchecker or wrote it while standing upside down. Why should I? The author put their name to the article and I got something out of it. That is all I was ever looking for.


You get it.


Not the author but it likely means running automated UI tests in the sim, yes. This involves running the app and programmatically selecting and sending interaction events.

Your previous experience was probably the agent running regular unit tests, which obviously don’t need ui environment, but mostly *do* need an iOS runtime, which is why it needs to boot the simulator.

An idiosyncrasy of the way unit tests are executed in Xcode is that they run from the actual app deployment target, and so while running unit tests you’ll also see any app initialisation and background tasks running at the same time. It’s quite a good idea to use compiler directives or launch arguments to disable the usual app setup in the App or App Delegate. Why this isn’t a built in option is beyond me, but it’s definitely confusing behaviour when you’re just running isolated tests!


Thanks! It's been over a decade since I've used Xcode/launched a native iOS app, so I wasn't sure what the capabilities were.

Looks like built-in UI testing was launched at WWDC 2015, so I missed it by a year!


Probably it’s less popular in America, but it’s huge in Asia, so I doubt the solitaire version is more well known globally


I only know the 4-player tile game, as a white dude from North America. But I only know from movies. I thought most people at least knew of it.


Yes, but most of HN is outside Asia, so I feel the clarification is helpful here.


Solitaire version should be pretty well known due to the computer game.


Thank you for this. Playing with my in-laws I’m always completely baffled by the scoring!


Do you have TestFlight installed? That’s the only other way I can think of delivering apps besides the App Store and MDM.


It’s amazing how many blank stares I get when I, as mobile engineer, tell stakeholders that we shouldn’t just implement some random interface idea they thought up in the shower and we instead need design input!

“But why can’t you just do it?” Because I recognise the importance of consistent UX and an IA that can actually be followed.

Just like developers, (proper) designers solve problems, an we need to stop asking them for faster bikes.


> “But why can’t you just do it?”

The answer should be "because users will hate it and use a competing product that's better designed".

A shame that it isn't actually true any more.


It should be, but it isn’t. Hence the reason “why can’t you just do what we asked” can often be followed by “then we will find someone who will” in the end.

Pushback is valuable until it becomes obstinance.

If we all somehow had their same crystal ball to know for certain that “stupid shower ideas” won’t work because a specific developer thinks they are bad, there wouldn’t be much need for R&D ever again. I suspect this developer doesn’t have one either, or I’d certainly like to buy it.


Those people generally are not eager for feedback especially if it's even remotely perceived as negative or some kind of gatekeeping.

The only way to avoid getting furious about this is to deeply understand that you can't require people to be properly self-aware, especially because many many people that checks the expert boxes are very incompetent or inadequate, so they when the come up with their half bake ideas, they delegate the other to deliver contrarian proofs. It's exhausting.


How about: It's too easy for users to use the product wrong, leading to unnecessary tech support calls, and doing tech support costs us money.


An underrated senior engineer skill is saving stakeholders from their own worst impulses.


Agreed, but I think it's underrated because the trait could also get someone fired or laid off.


Yeah you don't walk into the job doing it. But once your position is secure and you've built some stuff to where you can make changes in 5 minutes that might take someone else a day, then you can start pushing back on craziness.

There's also an art to how you do it. It's important to build up trust that you won't tell the client something is impossible or really hard just because you don't want to do it. You try to explain to them how much complexity this is going to add, how that will make it harder to add new features in the future, and most importantly offer some alternatives that get them 90% of what they're trying to do. Most clients will appreciate that approach IME, especially if you've already thrilled them a few times.

This is one of the reasons I'm not so worried about Claude taking my job in the immediate future. But I am still extremely worried about the industry as a whole and by extension the future of the middle class.


Yeah I agree. I've worked with my boss in various capacities for the last ten years. When he says "can we do X" my answer is always like "we can do anything, the question is does the company want to allocate Y resources to get X done." Claude, being a yes-man, will always say "you're absolutely right!" to any idea you throw at it, not knowing (and, perhaps more crucially, not caring) if it's the right fit for your product/business.

I think part of the problem is that many engineers don't stick around long enough to build that rapport, which isn't a problem of AI in itself but is certainly exacerbated by it.


Too many people think being good at designing a UI primarily means knowing where to put something on a page.


Honestly, if we could even get that level of thought, it would be an improvement.


There's a time and a place for it. If you already know exactly what the program needs to do, then sure, design a user interface. If you are still exploring the design space then it's better to try things out as quickly as possible even if the ui is rough.


The latter is an interesting mindset to advocate for. In almost every other engineering discipline, this would be frowned upon. I suspect wisdom could be gained by not discounting better forethought to be honest.

However, I really wonder how formula 1 teams manage their engineering concepts and driver UI/UX. They do some crazy experimental things, and they have high budgets, but they're often pulling off high-risk ideas on the very edge of feasibility. Every subtle iteration requires driver testing and feedback. I really wonder what processes they use to tie it all together. I suspect that they think about this quite diligently and dare I say even somewhat rigidly. I think it quite likely that the culture that led to the intense and detailed way they look at process for pit-stops and stuff carries over to the rest of their design processes, versioning, and iteration/testing.


Racing like in Formula 1 is extremely different from normal product design: each Formula 1 car has a user base of exactly 1: the driver that is going to use it. Not even the cars from the same team are identical for that reason. The driver can basically dictate the UX design because there is never any friction with other users.

Also, turnaround times from idea to final product can be insane at that level. These teams often have to accomplish in days what normally takes months. But they can pull it off by having every step of the design and manufacturing process in house.


There exist other ways to do the research. „Try things out“ is often not just a signal of „we don‘t know what to do“, but also a signal of „we have no idea how to properly measure the outcomes of things we try“.


I'm the lead for an internal tool for a non-technical team. We iterate so quickly that the team we're building it for was like "can you guys stop changing things so quickly? We can't keep up with where anything is." which was a fair assessment.


But that’s the point, no? Prototyping is useful but beyond a proof of concept, you still need a suitable user interface. I have no problems if there’s a rationale behind UI changes, but often we have stakeholders telling us to do something inconsistent just so their pet project can be presented to the user. That’s not design.


In my experience, it’s even more effort to get good code with an agent-when writing by hand, I fully understand the rationale for each line I write. With ai, I have to assess every clause and think about why it’s there. Even when code reviewing juniors, there’s a level of trust that they had a reason for including each line (assuming they’re not using ai too for a moment); that’s not at all my experience with Codex.

Last month I did the majority of my work through an agent, and while I did review its work, I’m now finding edge cases and bugs of the kind that I’d never have expected a human to introduce. Obviously it’s on me to better review its output, but the perceived gains of just throwing a quick bug ticket at the ai quickly disappear when you want to have a scalable project.


There is demand for non scalable, not committed to be maintained code where smaller issues can tolerated. This demand is currently underserved as coding is somewhat expensive and focused on critical functions.


What are some examples of when buggy code can be tolerated?


From recent personal examples

We have a somewhat complicated OpenSearch reindexing logic and we had some issue where it happened more regularly than it should. I vibecoded a dashboard visualizing in a graph exactly which index gets reindexed when and into what. Code works, a little rough around the edges. But it serves the purpose and saved me a ton of time

Another example, in an internal project we made a recent change where we need to send specific headers depending on the environment. Mostly GET endpoint where my workflow is checking the API through browser. The list of headers is long, but predetermined. I vibecoded an extension that lets you pick the header and allows me to work with my regular workflow, rather than Postman or cURL or whatever. A little buggy UI, but good enough. The whole team uses it

I'm not a frontend developer and either of these would take me a lot of time to do by hand


Leaf code. Anything that you won't have to build upon long-term and is not super mission critical. Data visualizers, dashboards, internal tools.

Pretty much everywhere where a 80% working tool is better than no tool, and without AI the opportunity cost to write the tool would be too high.


If the code is being used by a small group of people who are willing to figure out and share workarounds for those bugs - internal staff, for example.


Aren’t you also paying internal staff for their time. Waisting their time is waisting your money.


I've been in these situations before. If there's a known bug in an internal tool that would take the development team a day to investigate and fix - aka $10,000s - it's often smarter to send around an email saying "don't click the Froople button more than once, and if you do tell Benjamin and he'll fix it in the database for you".

Of course LLMs change that equation now because the fix might take a few minutes instead.


> If there's a known bug in an internal tool that would take the development team a day to investigate and fix - aka $10,000s - it's often smarter to send around an email saying "don't click the Froople button more than once, and if you do tell Benjamin and he'll fix it in the database for you".

How much will Benjamin's time responding to those calls cost in the long run?


Hopefully none, because your staff will read the email and not click the button more than once.

Or one of them will do it, Benjamin will glare at them and they'll learn not to do it again and warn their coworkers about it.

Or... Benjamin will spend a ton of time on this and use that to successfully argue for the bug to get fixed.

(Or your organization is dysfunctional and ends up wasting a ton of money on time that could have been saved if the development team had fixed the bug.)


> development team a day to investigate and fix - aka $10,000s

What about the non-fictional 99.999999999% of the world that doesn't make $1000/hour?


Large companies are often very bad at organizing work, to the tune of increasing the cost of everything by a large multiple over what you'd think it should be. Most of that cost wouldn't be productive developer time.


It costs them single digit thousands instead.


The alternative is the staff having no software at all to help with their task which wastes even more of their time.


You are setting up to say "I wouldn't tolerate that" for any example given, but if you look at the market and what makes people actually leave, instead of what makes people complain, then basically anything that isn't life-and-death, safety critical, big-money-losing, or data corrupting is tolerable. There's plenty of complaints about Microsoft, Apple, Gmail, Android, and all kinds of 3rd party niche business systems.

[Edit: DanLuu "one week of bugs": https://danluu.com/everything-is-broken/ ]

All the decades people tolerated blue-screens on Windows. All the software which regularly segfaulted years ago. The permeation of "have you tried turning it off and on again" into everyday life. The "ship sooner, patch later" culture. The refusal to use garbage collected or memory managed languages or formal verification over C/C++/etc because some bugs are more tolerable than the cost/effort/performance costs to change. Display and formatting bugs, e.g. glitches in video games. When error conditions aren't handled - code that crashes if you enter blank parameters. Bugs in utility code that doesn't run often like the installer.

One software I installed yesterday told me to disable some Windows services before the install, then the installer tried to start the services at the end of the install and couldn't, so it failed and finished without finishing installing everything. This reminded me that I knew about that, because that buggy behaviour has been there for years and I've tripped over it before; at least two major versions.

Another one I regularly update tells me to close its running processes before proceding with the install, but after it's got to that state, it won't let me procede and it has no way to refresh or rescan to detect the running process has finished. That's been there for years and several major versions as well.

One more famous example is """I'm not a real programmer. I throw together things until it works then I move on. The real programmers will say "Yeah it works but you’re leaking memory everywhere. Perhaps we should fix that." I’ll just restart Apache every 10 requests.""" - Rasmus Lerdorf, creator of PHP. I've a feeling that was admitted about 37 Signals and Basecamp, it was common to restart Ruby-on-Rails code frequently, but I can't find a source to back that up.


Points at the public sector


Any large enterprise. The software organisations write for themselves is pretty dire in most cases even without AI.


You need to have the AI write an increasingly detailed design and plan about what to code, assess the plan and revise it incrementally, then have it write code as planned and assess the code. You're essentially guiding the "Thinking" the AI would have to perform anyway. Yes, it takes more time and effort (though you could stop at a high-level plan and still do better than not planning at all), but it's way better than one-shotted vibe code.


The problem is those plans become huge. Now I have to review a huge plan and the comparatively short code change.


It shouldn't be any longer than the actual code, just have it write "easy pseudocode" and it's still something that you can audit and have it translate into actual coding.


TDD is a great way to start the plan, stubbing things it needs to achieve with E2E tests being the most important. You still need to read through them so it won't cheat, but the codebase will be much better off with them than without them.


This works but still lacks most context around previous tasks and it isn’t trivial to get it to take that into account.



I find it amazing that skills are essentially excellent tools for humans to understand too.


I really wish they were called lessons instead of skills. It makes way more sense and prevents the overloading of the term "skill".


There is some papers [0] showing that the skill and agent files reduce the reasoning effectiveness in some use cases (e.g. autogenerated)

[0] https://arxiv.org/abs/2602.11988

reference: https://news.ycombinator.com/item?id=47034087


We'd just be overloading "lessons" as well, and even more so because it takes more work to ground the concept, given its larger semantic distance from what we're describing.


> Even when code reviewing juniors, there’s a level of trust that they had a reason for including each line (assuming they’re not using ai too for a moment)

Even my seniors are just copy pasting out whatever Claude says. People are naturally lazy, even if they know what they’re doing they don’t want to expend the effort.


I hear you, but it seems quicker to predict whether the agent's solution is correct/sound before running it than to compose and "start" coding yourself. Understanding something that's already there seems like less effort. But I guess it highly depends on what you are doing and its level of complexity and how much you're offloading your authority and judgment.


If you click the full screen window, your little window is now behind it…


MacOS definitely has its issues but this just makes it sound like you have different expectations of how an OS should work. Different isn’t always bad. Hiding applications is a pretty key concept in MacOS. Shortcuts are pretty straightforward? Cmd+H to hide, Cmd+Q to quit. Spaces aren’t hidden- there’s lots of ways to access them, but it seems you haven’t bothered to learn them. In your example pressing ctrl+right would have switched the first full screen space. You could also have right clicked the Chrome icon in the dock for a list of windows.

BTW the dock doesn’t have to be hidden, and idk if it was a typo but alt+tab isn’t a default shortcut. Command is the key used for system shortcuts, so maybe you should have tried that? Like yeah it’s different but that doesn’t make it bad. If you been using it for 10 years without figuring that out…

—-

I’m with you on the 1st party apps though, and the stupid corners on Tahoe.


I call it "alt tab" because that's how my brain maps the keyboard. The reality is simple - I struggled going from Windows to Ubuntu about 20 years ago but ultimately made it to the other side knowing how to use both well. With macs, I didn't. 10 years later and all of my adaptations are to avoid the operating system. In 10 years the main thing I've learned is how to get myself out of a jam and stick to the parts of the OS that don't feel like shit. I mean, it's not like I haven't learned these things, I know how to gesture, I know how to exit full screen, etc, it's not like I didn't ever learn, I'm explaining that the experience was dog shit.

Anyone is free to claim that I just didn't try, or didn't give it a fair shake, or perhaps I'm just some idiot who doesn't know computers or whatever.

Maybe I just think an OS should work differently, but okay? I've never said that I have some sort of access to a platonic ideal of objective operating systems and that macs don't meet it. I'm saying that I think it's bad and I gave examples of why. And I think I can easily appeal to my experiences seeing others use the OS - I don't think they find anything you're talking about appealing either.


> Hiding applications is a pretty key concept in MacOS. Shortcuts are pretty straightforward? Cmd+H to hide, Cmd+Q to quit. Spaces aren’t hidden- there’s lots of ways to access them, but it seems you haven’t bothered to learn them.

They're not talking about Cmd+H hiding or virtual desktops - those exist on Windows too. The issue is how macOS handles window placement with zero visual feedback.

For example, when you open a new window from a fullscreen app, it just silently appears on another space. No indicator, no notification. You're left guessing whether it even opened and where it went. The placement depends on arcane rules about space layout, fullscreen ordering, and external displays - and it's basically random half the time. You either memorize the exact behavior or manually search through all your spaces.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: