That distinction is very close to how I’ve been approaching this in AIWebMastery.
For me, reproducibility is less about getting identical wording and more about being able to reproduce the same decision context and verify that it still passes the same checks.
I’ve found it useful to keep the model, prompt/version, tool inputs, captured tool results, decision, and outcome together. Then a replay can run in read-only mode against the captured context without automatically repeating side effects.
Write actions are treated separately and need an explicit approval before they are executed again. That also makes the audit trail much more useful: you can see what the agent observed, what it decided, what actually happened, and what was learned from the outcome.
So I’d say my target is primarily reproducible decisions + reproducible checks, rather than identical generated text. The wording can legitimately change when the model or context changes, while the decision should remain explainable against the same evidence.