NOTE 01
The model generates each shot from scratch
A video model does not remember the last shot unless you hand it something to go on. Each clip is a new generation, guided by whatever you give it: the text prompt and the reference images. If the only identity information is a written description such as "a woman in her thirties with short dark hair", every shot will interpret it a little differently, and the differences accumulate.
That is the first thing a reference image fixes. It replaces an ambiguous description with a concrete face, and it is why this workflow asks you to settle the cast before generating video.