How it works and why prompts are critically important here (unlike LLMs)
Recently, I announced my work on the OSS project OpenEffect (the repository is not published yet).
In short, OpenEffect is an open-source product designed to facilitate video generation from photos based on pre-prepared prompts.
The beta is on a temporary domain for my subscribers to check out examples and poke around the UI: OpenEffect
While working on our joint project OpenEffect, I encountered a non-obvious pain point. Prompting for video generation models is not as straightforward as it seems. Sometimes the model does not do what you have in mind.
We are used to LLMs (like ChatGPT) understanding us almost literally with half a word, and needing just a small unstructured prompt for the neural network to propose a solution to the problem and even give recommendations on related topics.
Why is it not so straightforward with video models?
Key Difference: LLM vs Video Models
LLM (text models)
- work with discrete tokens
- trained on structured language
- understand context, intent, and meaning well, even with inaccuracies
For example, if you ask the model: write a short post about ML
You will almost always get a satisfactory answer, even with short and inaccurate phrasing.
Video models (Kling, Wan)
- work with visual patterns
- trained on raw images and videos
- do not understand meaning but match: words to visual associations
In reality, the model does not understand what you want; it guesses what it looks like.
In an LLM prompt, you can write: “write beautiful code for an ML pipeline”
and the model will handle it.
When generating video, the prompt: “add dramatic flashes”
can yield different results: fireworks, flares, strange light artifacts
Example: applying a video effect to a photo paparazzi-flash
In short, the essence of the effect is to create short flashes from cameras, as if a person is being photographed by paparazzi.
input — a photo of a girl in a red dress
1.
Naive prompt:
Bright dramatic flash bursts around the subject
Result: sparks, bursts of light, strange light sources in the frame
Because the model interprets bright, bursts as any strong light effect, not a camera flash.
2.
Conducting tests, making it smarter:
clean bright white flashes
Result: the frame simply gets desaturated, sometimes the color of clothing and skin disappears
Because for “white flash,” the model creates a complete overexposure, not a localized flash.
3–4–5–10
We struggle a lot and eventually get a fairly decent result.
Resulting prompt
Medium close-up or medium shot of the same subject from the input image.
Brief paparazzi camera flashes pop from off-camera in one continuous shot,
producing short clean photographic light hits across the face, outfit, and background.
No visible fireworks, no sparks, no pyrotechnic light effects.
The same subject remains planted and recognizable throughout.
Only subtle blink, breathing, or slight hair movement is allowed.
Stylish celebrity photo-call energy, clean continuity, no cut, no scene replacement.
As a result: short flashes, without “explosions” and garbage, the face remains stable.
Why did this happen
-
I fixed the type of scene:
“Medium close-up or medium shot” -
I specified the physics of the effect, not just the idea:
“Brief paparazzi camera flashes pop from off-camera” -
I defined the form of light manifestation:
“short clean photographic light hits” -
I eliminated incorrect interpretations:
“No visible fireworks, no sparks, no pyrotechnic light effects”
and so on.
Therefore, the prompt works not because it is beautifully written, but because it sets all the main parameters for video generation.
The model simply chooses the most likely visual pattern.
If you write: “fire”
even in the context of the rest of the description, the model may interpret it as: an explosion, fire, burning something.
So it’s important to limit the space of options.
Example of prompt structure:
-
Basic action
“camera flashes appear” -
How exactly it happens
“short, clean, photographic” -
What limitations the model has
“no sparks, no explosions, no full-frame white wash”