> To change those percentages to 45% and 55% in a non learned way makes it seem like training wasn’t important?
I think (or I hope, anyway) that this overestimates how much impact the tweaks actually have on text.
The model has some things it "wants" to say. If it wants to tell a story about how someone reacted to dreary weather, it's going to tell approximately the same story regardless of whether the dice-roll caused it to describe the weather as "gray" or "overcast". And because "gray" and "overcast" were _already_ possibilities, the tweak from 52% -> 55% is completely lost in the noise.
But it's true that this is all based on hope. I'm confident that you could make the tweak against arbitrary prose and even a true artiste like Gruber would never be able to tell the difference. I'm less confident that there isn't some edge case somewhere that causes a tweak to be worse than 3%, especially in some narrow application where word choice _does_ matter (like law). Even then, though, laws are already written by people who are as noisy if not noisier than LLMs.
I think (or I hope, anyway) that this overestimates how much impact the tweaks actually have on text.
The model has some things it "wants" to say. If it wants to tell a story about how someone reacted to dreary weather, it's going to tell approximately the same story regardless of whether the dice-roll caused it to describe the weather as "gray" or "overcast". And because "gray" and "overcast" were _already_ possibilities, the tweak from 52% -> 55% is completely lost in the noise.
But it's true that this is all based on hope. I'm confident that you could make the tweak against arbitrary prose and even a true artiste like Gruber would never be able to tell the difference. I'm less confident that there isn't some edge case somewhere that causes a tweak to be worse than 3%, especially in some narrow application where word choice _does_ matter (like law). Even then, though, laws are already written by people who are as noisy if not noisier than LLMs.