Health “digital twins” can mislead unless built for cause-and-effect and change
This paper warns that copying engineering-style digital twins for use in health care can mislead clinicians and patients. The authors identify two traps—“fidelity” and “feedback”—that arise when a virtual patient model is judged only by how well it reproduces past data. They argue health digital twins should instead be designed as causally valid, modular, and evolving systems so their recommendations remain reliable as they are used.
The fidelity trap is the belief that a model that fits past patient data well will correctly answer “what if” questions about different treatments. The authors point out that predicting what happened is not the same as predicting what would have happened under a different decision. They use dynamic insulin dosing from continuous glucose monitoring (CGM) as an example: a model that reproduces observed glucose and dosing patterns may still fail to tell whether a larger insulin dose would improve outcomes, because past doses were chosen in response to rising glucose rather than causing it.
The feedback trap happens when the twin is updated on data that were shaped by its own recommendations. If the twin tells some patients they need less monitoring, fewer glucose readings will be recorded for that group. When the model is refit on this thinner data, it can appear more certain and reinforce the original recommendation—even if the real risk is unknown. The authors also note that recommended treatments are often not delivered exactly as suggested, because patients or clinicians adapt or override recommendations. These feedback loops can bias future estimates.
To avoid these problems the paper proposes three design principles. Causal validity means tying each recommendation to explicit assumptions about cause and effect, not just to predictive fit. Modularity means separating parts of the system that need interventional justification from parts used only for prediction. Governed evolution means updating the twin while tracking and adjusting for how its own use changes which patients are treated or measured. The paper develops this framework, illustrates it with the insulin example, and discusses implications for researchers, clinicians, system designers, and regulators.