While that might be ideal, is that really the case with most LLM training data? Does the curation process weed out all the slop from bad writers?
While that might be ideal, is that really the case with most LLM training data? Does the curation process weed out all the slop from bad writers?