8
My coworker said training AI on synthetic data was a bad idea and he was right
Tbh I thought feeding an LLM more clean generated data would fix its hallucinations. After 3 weeks of testing we got worse outputs that kept repeating the same fake facts. Now I'm back to scrubbing real datasets from public forums and it's actually working way better.
2 comments
Log in to join the discussion
Log In2 Comments
jamie_smith13d ago
Lol "repeating the same fake facts" - sounds like my brain on a Monday morning honestly. Guess we both learned the hard way that garbage in, garbage out applies even when the garbage is polished up fancy.
7
alice_kim13d ago
Are you saying that even when the AI talks smooth and sounds sure, it's still just repeating whatever garbage we fed it in the first place? Because that's what it feels like when I ask it for advice on a recipe and it gives me directions that would set my kitchen on fire. It's like the fancier the presentation, the harder it is to spot the nonsense underneath. What do you think makes people trust the polished version over the plain truth?
5