原來R2V先係H3既精髓
淨係俾1張圖同一段聲就已經得
原本諗住都係跟返First Frame咁去
但用Dasiwa Model V2 gen長片(15秒)會好易暴走(成個鏡頭轉晒)
用返原生model既話,雖然會跟返First Frame但有慢動作feel
重點係Dasiwa Model V2暴走得黎異常高質
keep到畫風之餘畫得仲豐富過我原圖
![[sosad]](/faces/sosad.gif)
(V1雖然唔會點暴走,但會嚴重少女化同畫風有變

)
所以直頭唔叫佢跟First Frame,直接畫個新畫面出黎
結果效果勁正 (冇試真人,2D真係超正)
Prompt跟返固有R2V格式
subject_definitions:
<Subject 1> is the women in <Picture 1>, 自己再寫一下囡囡特徵.
<Picture 1> is the character identity reference for <Subject 1>.
<Audio 1> is the voice-timbre reference for <Subject 1>.
summary:
[reference generation + audio reference] Use <Subject 1> from <Picture 1>, and the voice character of <Audio 1>.
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - identity remain consistent.
<Picture 1>: reference - the character identity are retained.
<Audio 1>: reference - timbre and delivery are followed without copying the signal.
detailed_description:
呢度填gemma 4 gen既prompt
overall_soundscape:
Anime girl style
non_diegetic_music:
N/A
Prompt (detailed_description)用gemma 4 gen
要用uncensored model先俾出咸野
https://huggingface.co/pekkAi/gemma-4-comfyui-text-encoder-heretic/tree/maingemma 4 gen text workflow參考
https://docs.comfy.org/tutorials/llm/gemma4/gemma4可以淨係餵圖俾條女佢做參考,叫佢設計場景、對白果的點寫
最後要附有example prompt俾佢跟[Shot 1]果的格式
出黎畫面真係有90分
聲就得5-6成似
而加可以半抽卡式直接gen自己想要既囡囡樣既唔同指定情景片了