You Can’t Tell Which LLMs Were Trained on What. Here’s Why That Matters.
The clean mental model of a ‘thin base LLM + RAG’ is a dangerous fantasy. Post-training and reinforcement learning now inject domain knowledge opaquely, making it impossible to separate language from facts. You can’t audit training data; you can only test outputs. Build rigorous evaluation instead of chasing provenance.