T
8

My coworker said training AI on synthetic data was a bad idea and he was right

Tbh I thought feeding an LLM more clean generated data would fix its hallucinations. After 3 weeks of testing we got worse outputs that kept repeating the same fake facts. Now I'm back to scrubbing real datasets from public forums and it's actually working way better.
2 comments

Log in to join the discussion

Log In
2 Comments
jamie_smith
Lol "repeating the same fake facts" - sounds like my brain on a Monday morning honestly. Guess we both learned the hard way that garbage in, garbage out applies even when the garbage is polished up fancy.
7
alice_kim
alice_kim13d ago
Are you saying that even when the AI talks smooth and sounds sure, it's still just repeating whatever garbage we fed it in the first place? Because that's what it feels like when I ask it for advice on a recipe and it gives me directions that would set my kitchen on fire. It's like the fancier the presentation, the harder it is to spot the nonsense underneath. What do you think makes people trust the polished version over the plain truth?
5