Hacker Newsnew | past | comments | ask | show | jobs | submit | graerg's commentslogin

Hey! Former astronomer here; I used to work a lot with SDSS data products. Do you happen to have a link to any sort of photometry DB/catalog for Rubin? Basically I’m looking to grab a ton of data but I’ve been out of the field for a while and I’m not sure where they publish their data sets and which ones you’d recommend as being relatively accessible.


You want to check out the early science release info at https://rubinobservatory.org/for-scientists/resources/early-... . This is all early science from preliminary data. It is mostly to help people ramp up developing their software against the type of things they will be seeing in the future. There some known issues and gotchas outlined in the release. Hopefully this will be enough to whet your appetite. The data seen in this sky viewer is part of the Data Preview 2.


I hope the data releases are a bit better than what I've seen from recent telescopes.

The common pattern seems to be that the majority of data is locked away behind multi-year embargo periods, or never released to the public at all with no explanation why, with only the smallest little crumbs of raw data actually released to the public, and usually years out of date. Also you have to go through some kafkaesque chain of technologies/servers/bureaucracy to get access to that little bit of data.

It's a little depressing to me when the public funds these telescopes and then get snubbed when it comes to getting access to the data they produce.


We are getting pretty far outside my area of expertise / knowledge on this topic here, but I know I and others agree with you here. There is a Rubin Observatory data access policy out there. Don't quote me on the exact specifics but I believe that most colleges and research institutions already have a mechanism to enable data access setup, or can get one. This is mostly an id auth type thing to make sure those accessing have data rights. There is a policy where people not associated with those can request access too, but I am not sure the procedure and would not be qualified to weigh in on anything more specific.


Curious what were your thoughts about this story when it broke: https://www.theatlantic.com/science/archive/2024/12/vera-rub...

And whether it impacts your work? For example, if the unnamed agencies' efforts leave a telltale mark in areas they change, that would still be problematic, so presumably they have to place noise corresponding to adjacent pixels. Are you concerned about how data integrity affects your work output? Penny to understand the internal scuttlebut / any internal grousings about this.


I built if-i-go-missing.com along these lines. Weird, I’m also a Brown!


Hahah, great minds think alike :)


Location: Orange County, CA (Pacific Time) Remote: Yes (preferred); hybrid Orange County also fine Willing to relocate: No Technologies: Python, Go, SQL | FastAPI, Django, Vue, HTMX | Postgres, TimescaleDB, DuckDB, Redis | Temporal, Airflow, NATS | AWS (S3/RDS/EC2/Redshift), Docker, Kubernetes | Prometheus/Grafana | LLMs in production (RAG, eval, orchestration, structured extraction) Resume/CV: https://brojonat.com/Jon_Brown_Resume.pdf Website / writing: https://brojonat.com GitHub: https://github.com/brojonat Email: brojonat@gmail.com

PhD-trained engineer/manager (astronomy by training -- data is data). Currently Senior Data Science Manager at Hyundai Motor America's North American Safety Office, where I build large-scale vehicle telemetry pipelines, LLM-driven classification and RAG over enterprise data, and led department-wide LLM/AI adoption from technical workshops to production SDLC.

Previously managed a data science team at HeadSpin: built data science pipelines, killed error-prone spreadsheet workflows with proper APIs and dashboards, reimplemented the flagship data product to reduce costs, latency, and churn.

Looking for senior data / platform sorts of roles. Equally happy IC or manager -- I just want to ship. Open to full-time or fractional/contract; happy to even work on small projects you want to send my way to help get to know each other.

I build a lot on my own time, and the side projects are probably the best read on how I work:

- forohtoo -- Go service for awaiting Solana payments via SSE/NATS, with a clean `Await(memo, amount)` SDK. Writeup: https://brojonat.com/posts/forohtoo/ - IncentivizeThis (https://incentivizethis.com) -- bounty-based authentic ad platform using LLM-as-judge across Reddit, YouTube, Twitch, Instagram, HN, TripAdvisor. Writeup: https://brojonat.com/posts/incentivizethis/ - If I Go Missing (https://if-i-go-missing.com) -- dead-man's-switch service (Go + Temporal + Postgres + Twilio + Stripe). - IYBI ("If You Build It [They Will Come]", https://iybi-twc.com) -- websites for independent workers, managed entirely by email. Customer pays via Stripe, emails an AI agent (Claude + Cloudflare Email Worker + R2 + K8s CronJob), the agent researches them, builds the site, and handles ongoing updates. My own consulting site runs on it.

Mostly Go or Python services running on my personal k8s cluster. Happy to talk shop or small projects we could work on together to assess fit.


I'm working on a competitive coding gameshow. I'm imagining a combination of great british bakeoff, battle bots, and dota. Basically contestants get dropped into a fully equipped dev machine (all the bells and whistles one could want/expect including neovim, agent harnesses, cool styling, etc and if you want you can always clone your dotfiles and stow them!). I've gotten a decent prototype that live streams from Fly.io sprites to twitch, and I'm able to voice over or have OpenAI do commentary on the match. I've got a demo here: https://www.twitch.tv/videos/2792893261. Still a ways to go, but it seemed like a fun way to tinker with Sprites.


How long do the contestants have?


For my demos I've been running 3 minute sessions but I'm planning to run 30-60 minute sessions depending on the scenario. I want to see people push the boundary of what can be done in an hour (with an agent or without, it's up to the contestant!) and ultimately have the match VODs serve as entertainment but also reference examples for "good" development workflows.


It is already very plausible (and has been since the 1950s) without the advent of LLMs. This is just another layer on top of the preexisting and very plausible existential threats we already face.


Detail it. Justify it.

Your comment about before LLMs is a non sequitur. Demonstrate that an LLM can kill everyone on the planet.


Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out.

There can be arms-races in domains that are unfathomable to the participants. A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses. Actors are not necessarily privy to understand the means by which they will lose, and humans have only existed in a small window of time in which we fashioned a manicured garden, in which that full understanding was briefly possible. It is not favoured in the universe for us to fully understand our environment imho

If the risk must be exhaustively detailed before it is given credence, we are already doomed, and deservedly so


>Task a squirrel with justifying the risk of a fox, but from the biomolecular level. That is the level of the task you are setting out.

Thats a really deep thought for a 12 year old.

>There can be arms-races in domains that are unfathomable to the participants.

You cant even justify LLMs as being unfathomable. Oh watch out I am fathoming them. You cant stop me fathoming all over the place.

>A small mammal will die a billion times over before it understands the evolutionary mechanisms and the genetic playing field on which it loses.Actors are not necessarily privy to understand the means by which they will lose, and humans have only existed in a small window of time in which we fashioned a manicured garden, in which that full understanding was briefly possible. It is not favoured in the universe for us to fully understand our environment imho

Non Sequitur. One that sounds like it was made up for that "What the Bleep" garbage.

>If the risk must be exhaustively detailed before it is given credence, we are already doomed, and deservedly so

The risk needs to be justified as something more substantial than weird people writing wannabe edgy messages on the internet. If someone on the internet told you that we need to drastically reverse living standards because there's a risk that modern technology will summon King Kong any reasonable person would ask for the working out instead of running for a cave.


You're kind of an asshole. No thanks


Its not like you handed me anything but woo to work with. There's really nothing less respectful than making up absolute nonsense and expecting a kind and thoughtful reply.


No they're right. Regardless if one agrees with you or not, doesn't change the fact that your behavior was that of an asshole. I would know since I'm one too.


Task a squirrel with agreeing with the behavior of a fox, but from the biomolecular level. That is the level of the asshole you are setting out.


> Thats a really deep thought for a 12 year old.

This was completely unnecessary. I understand why you say it, you like to make people feel bad. But it was being an asshole, regardless of how you try to justify it.

I'll invite to you our ethical sociopaths group if you want to join.


Location: Orange County, CA (Pacific Time) Remote: Yes (preferred); hybrid Orange County also fine Willing to relocate: No Technologies: Python, Go, SQL | FastAPI, Django, Vue, HTMX | Postgres, TimescaleDB, DuckDB, Redis | Temporal, Airflow, NATS | AWS (S3/RDS/EC2/Redshift), Docker, Kubernetes | Prometheus/Grafana | LLMs in production (RAG, eval, orchestration, structured extraction) Resume/CV: https://brojonat.com/Jon_Brown_Resume.pdf Website / writing: https://brojonat.com GitHub: https://github.com/brojonat Email: brojonat@gmail.com

PhD-trained engineer/manager (astronomy by training -- data is data). Currently Senior Data Science Manager at Hyundai Motor America's North American Safety Office, where I build large-scale vehicle telemetry pipelines, LLM-driven classification and RAG over enterprise data, and led department-wide LLM/AI adoption from technical workshops to production SDLC. Previously managed a data science team at HeadSpin: built data science pipelines, killed error-prone spreadsheet workflows with proper APIs and dashboards, reimplemented the flagship data product to reduce costs, latency, and churn.

Looking for senior data / platform sorts of roles. Equally happy IC or manager -- I just want to ship. Open to full-time or fractional/contract; happy to even work on small projects you want to send my way to help get to know each other.

I build a lot on my own time, and the side projects are probably the best read on how I work:

- forohtoo -- Go service for awaiting Solana payments via SSE/NATS, with a clean `Await(memo, amount)` SDK. Writeup: https://brojonat.com/posts/forohtoo/ - IncentivizeThis (https://incentivizethis.com) -- bounty-based authentic ad platform using LLM-as-judge across Reddit, YouTube, Twitch, Instagram, HN, TripAdvisor. Writeup: https://brojonat.com/posts/incentivizethis/ - If I Go Missing (https://if-i-go-missing.com) -- dead-man's-switch service (Go + Temporal + Postgres + Twilio + Stripe). - IYBI ("If You Build It [They Will Come]", https://iybi-twc.com) -- websites for independent workers, managed entirely by email. Customer pays via Stripe, emails an AI agent (Claude + Cloudflare Email Worker + R2 + K8s CronJob), the agent researches them, builds the site, and handles ongoing updates. My own consulting site runs on it.

Mostly Go or Python services running on my personal k8s cluster. Happy to talk shop or small projects we could work on together to assess fit.


That's the whole thesis; YAGNI.


I run my own temporal service in my k8s cluster; this setup is the backbone for almost all my applications. For simplicity, I opted for the postgres backend. You still need to run the 4 (?) other service (history, matching, frontend, ui, maybe others, definitely others if you want observability with prometheus/grafana, and tad bit more complexity if you want tailscale to get in there and poke around).

They ship Helm charts so reality is somewhere between "helm deploy" and "substantial ops burden". I don't have to touch it very frequently, but that is not to say I don't have to touch it. There's occasional releases and there have been times where (probably due to my inexperience with helm) I botched an upgrade and lost some data. And I've been on this journey for years; when I first started, they didn't have a Python SDK and it was one of my (many) excuses to learn Go. But anyway to your point, yes, if you're comfortable with k8s and Helm then you shouldn't have much of a problem running hundreds of thousands of workflows; if you want to really push the throughput and optimize cost you probably need to get creative the individual services and look into cassandra (maybe? idk).


They made a big point of explicitly advertising this as a feature with the GPT-5 rollout, no? Routing to cheaper models/less reasoning depending on the input prompt.


There's potentially some discussion of this publicly to investors. I feel there's more going on there and is re: quality issues described above.


Exactly! It must be exhausting to have this huge preoccupation with determining if something has come from an LLM or not. Just judge the content on it's own merits! Just because an LLM was involved doesn't mean the underlying ideas are devoid of value. Conversely, the fact that an LLM wasn't involved doesn't mean the content is worth your time of day. It's annoying to read AI slop, but if you're spending more effort suspiciously squinting at it for LLM sign versus assessing the content itself, then you're doing yourself a disservice IMO.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: