Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Because you need more training data for better results and they are running out of new training data.


I don't think so.

While it may be true that new data is coming in at a trickle these days, due to things like Discord, Slack, et al. all locking conversation and context up, as well as the daily volume of chapter is small relative to what is out there now.

The fact is that training data can be used in many different ways and I bet you we see the products of that fairly quickly as those who see this same as I do reach a point where they want to show n tell and test.


>The fact is that training data can be used in many different ways and I bet you we see the products of that fairly quickly as those who see this same as I do reach a point where they want to show n tell and test.

Sounds like wishful thinking to overcome the limitations of LLMs.

At the same time we get more and more texts generated by LLMs so it gets harder to get actual man made texts.


That’s true for LLMs but not necessarily for reinforcement learning




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: