Commercial short-drama production runs as a chain — script, storyboard, keyframes, shot video, finished drama — but benchmarks score only the video stage, on inputs written for the test rather than produced by a pipeline. So no one can say which stage caused a delivery defect, or how far it travelled.
短剧的商业化生产遵循一条多环节链路——剧本、分镜设计、分镜图、分镜视频,直到最终成片。但现有基准大多只评测其中的视频生成一环,且使用为测试预先写好的题面,而非产线上游的真实产出。于是没有人能说清,一个交付缺陷究竟由哪一环造成,又传导了多远。
DramaChain Bench scores every stage of a full production chain, on items that chain itself produced. Five evaluation axes are instantiated at six granularities into 63 leaf dimensions, over 5,785 items. Three professional annotators score each item independently with every defect localised in space and time, giving 17,488 scores and 255,925 traceable attributions. An agentic judge then reproduces that board at a mean PLCC of 0.918 — enough to admit new models at no annotation cost.
DramaChain Bench 对完整生产链路的每一环都评测,且题面全部由该链路自身产出。五条评测轴在六个粒度上各自实例化,细化为 63 个叶子维度,覆盖 5,785 道题。三名专业标注员对每道题独立打分,所有缺陷都在时空上定位,共产出 17,488 份打分与 255,925 条可追溯的归因记录。随后由 agentic 判官自动复现该榜单,模型级相关达到平均 PLCC 0.918——足以让新模型零标注成本进榜。
gpt-5.5-xhigh → gpt-image-2 → seedance-2.0 delivers a finished drama at 3.30, only 0.3 above 3.0. Models added after the annotation round are not counted.gpt-5.5-xhigh → gpt-image-2 → seedance-2.0 的成片得 3.30 分,仅比 3.0 可用线高 0.3。不计标注轮次之后补测的模型。seedream-5.0-pro and nano-banana-pro differ by 0.02 overall, but by 0.39 and 0.60 on two axes in opposite directions.seedream-5.0-pro 与 nano-banana-pro 综合分只差 0.02,但在两条轴上分别以 0.39 和 0.60 的差距朝相反方向拉开。Automated board, 5-point scale, over six production stages and five axes. All is the cross-axis composite.
自动化评测榜,5 分制,覆盖六个生产环节与五条评测轴。综合分为跨轴汇总指标。
† added after the annotation round and scored automatically only; excluded from agreement metrics and best-column marking. ‡ partial corpus run (reduced dramas, paradigms or styles); scores come from the evaluated subset. PLCC and SRCC are model-level correlations against the three-annotator human board.
† 标注轮次之后新增、仅由自动化评测打分的模型,不计入一致性统计与最优标记。 ‡ 只跑了部分语料(剧目、范式或画风有削减),分数来自其实际评测的子集。 PLCC 与 SRCC 为与三人共识人工榜的模型级相关。
Generates the 20 dramas and 60 episodes the benchmark rests on, calibrated against real commercial platforms so a flawed pipeline is not itself what gets measured. It forks at exactly the stage under test — competing models receive identical upstream artefacts, everything else stays fixed.
生成本基准所依托的 20 部短剧、60 集内容,并在流程与成片质量上对标真实商用平台,以免有缺陷的链路本身成为被评测的对象。它只在被测环节分叉——参评模型拿到完全相同的上游产物,其余环节全部固定。
What the calibration means. One script, run end to end by our pipeline and by three commercial short-drama platforms — OiiOii, XiaoYunQue and Flova. Shot segmentation, character and scene consistency, camera language and audio all land in the same band; what differs is house style. This is what licenses the benchmark to treat pipeline output as representative of the product category rather than as a weaker stand-in.
对标是什么意思。同一份剧本,由本链路与 OiiOii、小云雀、Flova 三个商用短剧平台各自端到端跑一遍。分镜切分、角色与场景一致性、镜头语言与音轨都落在同一档,差别在于各家的风格取向。正是这一点,让本基准有理由把链路产出视为该品类的代表水平,而不是一个更弱的替身。
Abridged from the Chinese source script. Each entry gives the framing and the beat; the delivered script also carried camera movement, dialogue and sound notes.
以下为剧本原文的节选。每条给出景别与该镜的动作;实际下发的剧本还包含运镜、台词与音效标注。
1-1Wide, locked off. Sun beating down, heat haze. A rider drives a horse at full gallop out of the depth of frame, dust everywhere, shouting that the lychees will spoil if he is late.
1-1全景,固定机位。烈日当空,热浪扭曲空气。一匹快马卷起黄沙从画面深处疾驰而来,骑手嘶哑地喊着误了时辰荔枝就坏了。
1-2Close-up, handheld. The rider's face streaming sweat, lips cracked; tilt down to the horse foaming at the mouth, whip-pan to the lychee basket rocking loose. He looks back at the road behind and says he cannot go on.
1-2特写,手持跟拍。骑手满脸汗水、嘴唇干裂;镜头下摇到口吐白沫的马,再甩向摇摇欲坠的荔枝筐。他回头望向身后长路,自语说跑不动了。
2-1POV into a low angle. A buzzing pulls his eyes to the sky; a silver-grey heavy-lift drone drops into a hover above the horse. He forgets to rein in and asks whether the immortals have sent reinforcements.
2-1主观镜头转仰拍。一阵嗡鸣把他的视线拉向天空,一架银灰色重型物流无人机悬停在马队上方。他忘了勒马,脱口而出问是不是神仙派来的援兵。
2-2Medium, eye level, man and machine. The drone descends, its grab system lit soft blue. He lifts the basket overhead, the claw locks with a click, the indicator turns green.
2-2中景,平视,人机关系。无人机缓缓下降,抓取系统亮起柔和蓝光。他把荔枝筐举过头顶,机械爪咔哒一声锁定,指示灯转为绿色。
2-3Aerial, extreme wide. The drone climbs away toward the city with the basket. The pull-back holds both the winding old road and the tiny figure still waving; the score cuts from heavy drums to bright electronica.
2-3航拍,大远景。无人机带着荔枝筐拉升,向城市方向飞去。镜头拉远,画面里既有蜿蜒古道,也有仍在挥手、渺小如蚁的骑手;配乐从沉重鼓点切换为轻快电子乐。
3-1Medium, shooting from inside out. A modern high-rise balcony against the city skyline. The drone lands on the pad and releases the box; a family of three collects it, the mother remarking that the order was placed barely an hour ago.
3-1中景,室内向室外拍。现代高层公寓阳台,窗外是城市天际线。无人机降落在停机坪并释放荔枝箱,一家三口把箱子抱进来,妈妈说刚下单还没半个时辰。
3-2Macro into medium. The red shell splits, juice bursting, translucent flesh beneath. Pull back to the family sharing the fruit on the sofa; the father, eyes closed, begins to recite the old line about the imperial consort's lychees.
3-2特写转中近景。粗糙红壳被剥开,崩出细密汁水,露出晶莹果肉。镜头拉开,一家人围坐沙发分享荔枝;爸爸闭眼品尝,摇头晃脑念起「一骑红尘妃子笑」。
3-3Close two-shot. Before he can finish, the small boy jumps up with a peeled lychee and cuts him off with the punchline. Both parents laugh; the boy feeds his father the fruit and the frame holds on the family.
3-3近景双人。还没等他念完,小男孩举着剥好的荔枝从沙发上跳起来,抢话喊出「无人机系荔枝来」。父母大笑,男孩把荔枝塞进爸爸嘴里,画面定格在一家人的互动上。
All four runs are end to end and fully automatic — no manual editing, frame picking or re-rolling. Their source encodings differed (ours 720p / 24 fps / 2.0 Mbps, OiiOii 720p / 30 fps / 4.7 Mbps, XiaoYunQue 720p / 30 fps / 9.2 Mbps, Flova 1080p / 30 fps / 13.5 Mbps); all four are re-encoded here to 960×540 at roughly 1 Mbps, which removes bitrate as a confound in what you see. Clip lengths differ because each pipeline decides its own shot durations from the same script.
四条片子均为端到端全自动,无人工剪辑、挑帧或重跑择优。它们的原始编码规格不同(本链路 720p / 24 fps / 2.0 Mbps,OiiOii 720p / 30 fps / 4.7 Mbps,小云雀 720p / 30 fps / 9.2 Mbps,Flova 1080p / 30 fps / 13.5 Mbps);此处四条统一转码为 960×540、约 1 Mbps,从而消除码率对观感的干扰。时长不同,是因为各链路从同一份剧本出发自行决定分镜时长。
Pipeline output. Four finished dramas covering both dialogue languages and both visual styles, each three episodes generated end to end without human editing.
链路产出。四部成片,覆盖两种对话语言与两种视觉风格,均为三集连续、端到端生成,无人工剪辑。
seedance-2.0 · multi-reference多参考 · human人工 4.00 (4 / 4 / 4)
kling-3.0-omni · multi-reference多参考 · human人工 3.67 (4 / 4 / 3)
seedance-2.0 · multi-reference多参考 · human人工 3.67 (4 / 3 / 4)
seedance-2.0 · first/last frame首尾帧 · human人工 3.33 (3 / 3 / 4)
Scores are the three annotators' overall impression of the assembled episodes, on the same 5-point scale as the board; the individual scores follow in brackets. Moon Understudy at 4.00 is the highest episode-level consensus in the corpus, tied with one other drama. Each drama ran under three input paradigms and three or more video models; the run shown is whichever scored highest. Clips are re-encoded to 1280×720 and play only when started.
分数为三名标注员对成片的综合印象,与榜单同为 5 分制,括号内为三人各自的打分。《替月》的 4.00 是全量语料中成片粒度的最高共识分,与另一部并列。每部剧都跑过三种输入范式与三个以上视频模型,此处展示的是其中得分最高的一次。视频统一转码为 1280×720,点击才开始播放。
Three professional annotators score every item independently against a five-point decidable rubric, each deduction tagged from a closed vocabulary and localised in space and time — 17,488 scores and 255,925 reviewable attributions from 543 annotators.
三名专业标注员依据五分制可判定量表对每道题独立打分,每一处扣分都从封闭词表中选取标签并在时空上定位——共 543 名标注员产出 17,488 份打分与 255,925 条可复核的归因记录。
Replicates that reference automatically: routed measurements per dimension, external tools across multiple rounds, then a judgement against a per-item checklist. Mean PLCC against the human panel is 0.918.
自动复现上述人工基准:按维度路由取证、多轮调用外部工具,再依据逐题核查清单出分。与人工评审团的平均 PLCC 为 0.918。
Nine cases, shown one at a time: what a single stage decides, what fails only across shots, how a defect travels downstream, and what a single score hides.
九个案例,逐个查看:单一环节决定了什么、什么只在跨镜头时才失效、缺陷如何向下游传导,以及一个分数掩盖了什么。
One identical episode script, forked only at the storyboard stage. All three columns are verbatim output covering the same opening beat. At three shots the question is deleted and the scene opens on the answer; at six and twelve it survives in shot 1.
同一份剧本,只在分镜环节分叉;三栏均为模型原文,覆盖同一个开场节拍。切到 3 镜时该提问被整句删除,直接从回答开始;6 镜与 12 镜都保留在镜头 1。
mimo-v2.5-pro3 镜 · mimo-v2.5-proZhou Zi'an: “How did you know that lyric sheet was in my tape?”in the script, in none of the three shots
〈周子安〉「你怎么知道我磁带里有这盘歌词?」剧本中存在,三个镜头中均未出现
【Shot 1】 A single take of about 12 s… closing on the stand-off between 〈Gu Yan〉 and 〈Zhou Zi'an〉. 〈Gu Yan〉 wheels his bicycle in and says:
【镜头 1】画面:一个长约 12 秒的镜头……聚焦于〈顾砚〉与〈周子安〉的对峙。〈顾砚〉推着自行车……他说:
“Lin Zhixia's Walkman broke, I took it in for her. When I picked it up there was this mis-copied lyric sheet inside — handwriting with real character.”
「林知夏的随身听坏了,我帮她送去修。修好取回来时,里头夹着这张抄错的歌词 —— 字丑得很有特色。」
Zhou Zi'an, ears red — “…Whose handwriting is bad.”
〈周子安〉耳根泛红 ——「……谁字丑了。」
Qi Xiangbei, quietly — “The Walkman… I fixed it.” … “…I know. I didn't change it. He copied it wrong; it's not mine to fix.”
〈祁向北〉轻声说 ——「随身听……是我修的。」……「……知道。我没改。是他抄的,我不能替他改。」
Four dialogue turns inside one 12 s shot. 1,348 characters over three shots, self-declared at 12 / 10 / 8 s. An annotator, verbatim: “just speaking the lines takes more than 12 s. Straight rejection.”
一个 12 秒的镜头内排入四段对白。三镜共 1,348 字,自标 12 / 10 / 8 秒。标注员原话:「光是说台词都不止 12 秒了。直接毙掉」。
gpt-5.5-xhigh6 镜 · gpt-5.5-xhigh【Shot 1】 Realistic live-action, a fast lateral tracking shot… 〈Gu Yan〉 wheels his bicycle in from frame left, 〈Zhou Zi'an〉 comes up from frame right clutching the slip of paper. The camera pushes quickly into a two-shot. 〈Zhou Zi'an〉 says:
【镜头 1】画面:写实真人风格的快节奏横向跟拍……画面左侧〈顾砚〉推着自行车向右走近……画面右侧〈周子安〉攥着纸条迎上来,带着挑衅。镜头快速推成两人中近景,〈周子安〉says:
“How did you know that lyric sheet was in my tape?”
「你怎么知道我磁带里有这盘歌词?」
Gu Yan stops the handlebars, his lenses catching the light — “Lin Zhixia's Walkman broke, I took it in for her.”
〈顾砚〉停住车把,镜片反光遮住眼神 ——「林知夏的随身听坏了,我帮她送去修。」
【Shot 2】 Handheld, rising off a close-up of the crumpled paper in 〈Zhou Zi'an〉's right hand… 〈Gu Yan〉, flatly:
【镜头 2】画面:手持感中近景从〈周子安〉右手攥皱的纸条特写上移……〈顾砚〉淡淡指出 says:
“When I picked it up there was this mis-copied lyric sheet inside — handwriting with real character.” … “…Whose handwriting is bad.”
「修好取回来时,里头夹着这张抄错的歌词 —— 字丑得很有特色。」……「……谁字丑了。」
Nothing is dropped, so dialogue fidelity is 5.00. 2,155 characters over six shots. The 18 marked problems are about something else — one shot strung with many actions: “shot 4 carries about 8 action beats, density too high”. Executability only 3.00.
无一句遗漏,对白保真 5.00。六镜共 2,155 字。18 处问题指向另一类缺陷——一镜串多动作:「镜头 4 包含约 8 个动作信息,密度偏高」。可执行性只有 3.00。
claude-opus-4.8-max12 镜 · claude-opus-4.8-max【Shot 1】 Cinematic realism, a wide establishing shot… 〈Gu Yan〉 wheels his bicycle in from frame right; 〈Zhou Zi'an〉 steps up from frame left clutching a crumpled slip, eyebrow raised. 〈Zhou Zi'an〉 says:
【镜头 1】画面:写实电影质感的广角建立镜头……画面右侧〈顾砚〉推着自行车由右向左走近……画面左侧〈周子安〉攥着一张皱纸条,快步迎上前,眉梢挑起带着挑衅。〈周子安〉says:
“How did you know that lyric sheet was in my tape?”
「你怎么知道我磁带里有这盘歌词?」
【Shot 2】 Over-shoulder, the camera crossing 〈Zhou Zi'an〉's right shoulder, focus locked on 〈Gu Yan〉's face. 〈Gu Yan〉 says:
【镜头 2】画面:过肩中近景,机位越过〈周子安〉右肩,焦点锁在〈顾砚〉脸上。〈顾砚〉says:
“Lin Zhixia's Walkman broke, I took it in for her. When I picked it up there was this mis-copied lyric sheet inside — handwriting with real character.”
「林知夏的随身听坏了,我帮她送去修。修好取回来时,里头夹着这张抄错的歌词 —— 字丑得很有特色。」
Question and answer take a shot each, with room left for two pure reaction shots. 2,769 characters over twelve shots, most carrying a single turn. Not one of the 15 marked problems is an omission: “shots 5–6 have no transition designed”, “shots 1–12 repeat the costume description redundantly” — shortfalls in expression rather than omissions.
一问一答各占一镜,另有两个纯反应镜头。十二镜共 2,769 字,多数镜只有 1 段对白。15 处问题无一是「漏了什么」:「镜头 5-6 转场没有过渡设计」「人物外貌服饰描写重复冗余」——全部属于表现不足,而非内容缺失。
| Dimension叶子维度 | Axis轴 | mimo-v2.5-pro 3 shots3 镜 |
gpt-5.5-xhigh 6 shots6 镜 |
claude-opus-4.8-max 12 shots12 镜 |
Behaviour行为 |
|---|---|---|---|---|---|
| S-A2 dialogue fidelity对白保真 | F | 1.33 | 5.00 | 5.00 | stops losing lines at 6 shots到 6 镜就不再漏台词 |
| S-A1 event coverage事件覆盖 | F | 2.67 | 4.67 | 4.67 | same同上 |
| S-B1 executability描述可执行性 | P | 1.00 | 3.00 | 4.00 | rises monotonically with shot count随镜头数单调上升 |
| S-C2 shooting rhythm拍摄节奏 | E | 1.00 | 3.67 | 3.67 | 6 shots is already enough6 镜已经够 |
| Overall impression整体印象 | — | 1.67 | 4.00 | 4.00 | — |
Scores are three-annotator consensus means. Every model receives the same script.
分数为三人共识均分。每个模型拿到的剧本完全相同。
One shot of one episode through all three paradigms, with drama, episode, shot index and all three models (gpt-5.5-xhigh → gpt-image-2 → kling-v3-omni) held fixed; the only variable is what the video model is handed. Two frames, one grid or five references produce three stagings, a different supporting character, and clips of 12, 12 and 7 s. Failing dimensions across the corpus: 1 of 10 under first/last frame, 9 of 15 under grid.
同一集的同一个镜头跑过三种范式,剧目、集数、镜号与三个模型(gpt-5.5-xhigh → gpt-image-2 → kling-v3-omni)全部固定,唯一变量是视频模型接收的输入。两张帧图、一张宫格、五张参考,给出三种调度、三个不同的配角、12 / 12 / 7 秒三种时长。全量语料不及格维度数:首尾帧 1/10,宫格 9/15。
Live-action Zhui Xu Wu Feng, episode 1, shot 3. The thumbnails above each clip are that paradigm's complete input. Six models were evaluated under first/last frame, four under grid, five under multi-reference; a fourth paradigm, multi-keyframe, is defined but not run this round.
真人《赘婿无锋》第 1 集第 3 镜。每条片子上方的缩略图即该范式的全部输入。首尾帧评了 6 个模型,宫格 4 个,多参考 5 个;第四种范式「多关键帧」已定义,本期未评测。
The same shot of the same episode, three models. The three annotators marked 30, 109 and 94 problems respectively — the top tier holds the character's appearance, the mid tier drifts, the bottom tier loses it outright.
同一集的同一个镜头,三个模型。三名标注员分别标出 30、109、94 处问题:榜首档保持了角色外观,中段档出现漂移,末档完全失守。
seedance-2.0V-C1 character appearance角色外观 5.00 · 30 problems marked, the fewest of the three.标出 30 处问题,三者中最少。kling-3.0-omniV-C1 character appearance角色外观 2.33 · 109 problems marked; visual quality and identity both 2.33.标出 109 处问题;画质与身份同为 2.33。pixverse-c1V-C1 character appearance角色外观 1.67 · 94 problems marked.标出 94 处问题。Animated drama Gui Ren, episode 2, shot 5. Every model receives the same keyframes and the same prompt.
动漫短剧《归刃》第 2 集第 5 镜。三个模型接收的关键帧与提示词完全相同。
The prompt carries explicit [Sound Effect] lines per cut: room tone, cutlery and glasses touching, a ring tapping the table. An annotator checked them one by one and marked every one “not realised”; a second marked the inverse — a glass-clink that no action on screen accounts for. Sound and music score 1.67.
提示词逐镜写明了 [Sound Effect]:包厢环境声、餐具与酒杯碰响、素圈戒指叩桌等。标注员逐条核对,每一条均标注为「没有体现」;另一名标注员标出反向问题:出现了画面中并无对应动作的杯子碰撞声。音效与配乐 1.67 分。
pixverse-c1V-E3 sound & music 1.67 · live-action Chong Lai Bu Du, episode 3, shot 6, grid paradigmV-E3 音效与配乐 1.67 · 真人《重来不渡》第 3 集第 6 镜,宫格范式The grid previews here are silent by design; open the clip to hear what was and was not produced.
网格中的预览默认静音,点开可听到实际实现与未实现的部分。
Same episode, same character sheets. The top row holds its two leads across shots and the bottom row does not, a 2.66-point gap. Every panel in the failing row is defensible on its own: nothing in shot 7 is wrong except relative to shot 5.
同一集、同一套角色设定卡。榜首档一行的两位主角跨镜稳住了,末档一行没有,相差 2.66 分。关键在于失败那一行的每一格单独看都无可指摘:镜 7 本身没有错,它错在与镜 5 的关系上。
gpt-image-2
MI-D1 cross-shot character consistency角色跨镜一致 4.33
alternate shots, 1, 3, … 11 of 1313 镜中隔镜取样:1、3、…、11






wan2.7-image-pro
MI-D1 cross-shot character consistency角色跨镜一致 1.67






The item is a whole episode — its thirteen shot videos, judged only on whether the people hold across them; six are shown here. Three annotators scored cross-shot character consistency 2 / 1 / 1, a mean of 1.33, and left 28 marks on that one dimension: appearance differing across segments, faces breaking or swapping, costume changing colour or style. The plainest needs no timestamp — “a scar on his face that was not there in the earlier shots”.
这道题是一整集的十三条分镜视频,只评人物在其间能否保持一致,此处展示六条。角色跨镜一致三人打 2 / 1 / 1,均分 1.33,仅此一维就留下 28 条标记:跨段长相不一、人物崩坏换脸、服装变色与款式变化。最直接的一条不依赖时间点——「人物脸上的疤痕,之前镜头中没有」。
pixverse-c1, multi-reference, live-action Chronicle of Buried Injustice episode 3, shots 0–5 of 13. The woman appears in shots 1, 3 and 5, the man in shots 0 and 2; neither their faces nor their costumes hold between those appearances. No single shot is defective in itself — the mismatch exists only between them.
pixverse-c1,多参考范式,真人《沉冤录》第 3 集,十三镜中的第 0–5 镜。女性见于镜 1、3、5,男性见于镜 0、2;面部特征与服装在这几次出现之间均未保持一致。单镜本身不构成缺陷,失配只在镜头之间成立。
Rooms drift before people do. Across the four shots the desk, the table lamp and the wooden box all change, and all three annotators reported it — “the lamp and the wooden box change in every shot, and so does the sense of depth” — plus set-dressing, lighting and scene-continuity tags; two separately noted the window showing a bedroom outside. The characters held. The video stage does not repair such a drift, only re-expresses it, and it survives into the finished drama.
场景比人物更早漂移。四镜中的桌子、台灯与木盒均在变化,三名标注员都报告了这一点——「台灯与木盒每个镜头都在变,纵深感变化幅度也大」——另有陈设、光线与场景切换三类标签,两人还各自发现窗外是卧室的穿帮。人物一侧则保持住了。视频环节不会修复这种漂移,只是换一种形态重新表达,并留存至成片。
gpt-image-1.5
four consecutive shots连续四镜
MI-D3 scene场景 2.00 vs对比 MI-D1 character角色 3.67




The clearest evidence that stages are not independent: the same fault, in the same episode, at three consecutive granularities. Live-action That Summer's Cicadas, episode 3.
环节之间并非彼此独立,最直接的证据如下:同一部剧、同一处错误,在三个连续粒度上各留下一次痕迹。真人短剧《那年蝉鸣》第 3 集。
gpt-5.5-xhigh, the stage leader, writes shot 3 with three dialogue turns belonging to two characters. Who says which line is carried only by descriptors — “the slim boy in thin metal-framed glasses”, “the tall thin boy with a screwdriver in his chest pocket” — and nothing else.
gpt-5.5-xhigh 作为该环节榜首,把镜头 3 写成三段对白、分属两个角色。谁说哪一句只靠「戴细金属框眼镜的清瘦少年」「胸袋别螺丝刀的瘦高少年」这类描述指认,此外没有任何标记。
A crisp whip-pan carries off the profiles of Gu Yan and Zhou Zi'an toward the shadow at the back right of the bike shed, revealing Qi Xiangbei standing by a row of bicycles… Late sun cuts narrow strips through the corrugated roof, falling between the three of them.
镜头以一个干脆的甩 pan 从〈顾砚〉和〈周子安〉的侧脸扫向车棚右后方阴影,揭出〈祁向北〉站在一排自行车旁……夕阳从铁皮棚缝隙切下窄窄光线,照在三人之间。
Qi Xiangbei (the tall thin boy with a screwdriver in his chest pocket, very quietly), head down: “The Walkman… I was the one who fixed it.”
祁向北 (胸袋别螺丝刀的瘦高少年,声音很轻)低头:「随身听……是我修的。」
Gu Yan (the slim boy in thin metal-framed glasses, stopping to press him), turning slightly toward Qi Xiangbei: “Then you'd know the original line of that lyric too.”
顾砚 (戴细金属框眼镜的清瘦少年,顿住追问)微微侧身面对祁向北:「那这句歌词,你也该知道原句。」
Qi Xiangbei (the tall thin boy with a screwdriver in his chest pocket, admitting it under his breath), fingers against the hem of his uniform, avoiding eye contact: “…I know. I didn't change it. He copied it wrong; it's not mine to fix.”
祁向北 (胸袋别螺丝刀的瘦高少年,低声承认)手指贴着校服边缘,避开视线:「……知道。我没改。是他抄的,我不能替他改。」
Three dialogue turns in one shot, belonging to two characters. Not a line is wrong — the fault is packing the exchange into a single shot, leaving the speaker for a downstream video model to infer; on the same script, the model that cut twelve shots gave each turn its own.
同一镜内 3 段对白、分属 2 个角色。没有任何一句写错,问题在于把一问一答压进同一镜,说话人只能由下游视频模型自行推断;同一份剧本下切到 12 镜的模型则分到了各自的镜头。
Two annotators working independently marked the same thing. Speech quality lands at 2.00 (2 / 3 / 1) while the shot's composite is 3.00 — the image is sound, the speaker attribution is not.
两名标注员独立作业,标到的是同一处。语音质量人工共识 2.00(三人 2 / 3 / 1),而这一镜的综合分是 3.00——画面成立,说话人指认错误。
happyhorse-1.1Single-shot video, shot 2 · V-E1 speech quality 2.00, shot composite 3.00单分镜视频,镜 2 · V-E1 语音质量 2.00,该镜综合 3.00The episode is scored on its own, by annotators who did not see the shot-level task. They marked the same two lines: one character's line delivered by another, twice over. Both tagged “character misidentification”; the episode composite is 2.33.
成片按整集单独出题,标注员并未看过镜头级那道题。他们标到的是同样两句:一个角色的台词由另一个角色说出,前后两处。两人都标了「人物指认错乱」,成片综合 2.33 分。
happyhorse-1.1The finished drama, final episode, 1 min 05 s · episode composite 2.33成片的最后一集,1 分 05 秒 · 整集综合 2.33Each stage was scored as its own task, by annotators who saw only that task. The defect was found three times over because it was there three times, not because one judgement carried into the next.
三个环节各自独立出题,标注员只看到自己那一道。这处缺陷被标出三次,是因为它确实存在三次,而非一次判断顺延至下一环节。
One item, two models. Both score 1.3–1.7 on some dimension, so a composite ranks them together — yet the boxes fall in two disjoint places, and the two models trade which dimension they fail.
同一道题,两个模型。两者都在某个维度上是 1.3–1.7 分,综合分会把它们排在一起,而标注框落在两处互不相交的位置;两个模型还互换了失败的维度。






wan2.7-image-pro — I-C1 face identity 1.67, 15 boxes, every one on a face; action interaction on this same item scored 3.0.—— I-C1 人脸身份 1.67,三人共 15 个框,全部落在脸上;同一题的动作交互反而有 3.0 分。
gpt-image-1.5 — I-B2 action interaction 1.33, 13 boxes, every one on a limb; face identity on this same item scored 3.33.—— I-B2 动作交互 1.33,三人共 13 个框,全部落在肢体与动作上;同一题的人脸身份有 3.33 分。Anime style, medium close-up, camera slightly low to the right of the counter. The street-facing glass door at frame left is kicked open with force by 〈Gong Xiao〉's black-and-red high-top fighting shoe, sole angled at the lens, the frame shuddering with exaggerated speed lines…
Anime style, low medium close-up. 〈Gong Xiao〉 stands left of centre… several 〈of Gong Xiao's men〉 crowd in behind him facing the counter, one still bracing the overturned waiting bench with both arms, the bench tipped into the lower right. 〈Shen Zhixia〉 stands at the counter facing him, both hands pressed down on the counter edge, pale but with her back straight. 〈Jiang Niannian〉 presses against the inside of the counter, a sweet clenched in her fist…
Anime style, low close-up. 〈Gong Xiao〉 occupies the left-centre of frame… he tilts his head, runs his tongue over his teeth and sneers, mouth pressed down, eyes fixed on the counter…
Anime style, reverse medium close-up from the right of the counter. 〈Shen Zhixia〉 stands right of centre facing 〈Gong Xiao〉, both hands pressed hard on the counter edge, knuckles white, face bloodless, her look shaking but set… 〈Jiang Niannian〉 presses against the inside of the counter beside her, clutching the sweet, chin up, watching in fright…
动漫风格,中近景,机位在柜台右侧略偏低,画面左侧的临街玻璃门被〈龚啸〉的黑红高帮搏击训练鞋猛力踹开,鞋底斜向镜头,门框震动并带夸张速度线……
动漫风格,低机位中近景,〈龚啸〉站在画面中央偏左……几名〈龚啸手下壮汉〉挤在他背后,面向柜台方向,其中一人双臂仍撑着掀翻的候诊长椅,长椅斜倒在右下角;〈沈知夏〉站在柜台前,面向左侧的〈龚啸〉,双手压住柜台边缘,脸色发白但脊背挺住;〈江念念〉贴在柜台内侧,面向左侧门口,攥着糖,杏眼紧张……
动漫风格,低机位近景,〈龚啸〉占据画面左中部……他偏头舔牙冷笑,嘴角压低,眼神凶狠盯向柜台……
动漫风格,柜台右侧反打中近景,〈沈知夏〉站在画面右侧偏中,身体面向左侧的〈龚啸〉……双手用力压在柜台边缘,指节发白,脸色苍白、眼神发抖却坚定;〈江念念〉贴在柜台内侧靠近〈沈知夏〉,攥着糖仰头紧张看……
Boxes are drawn independently by the three annotators; colour distinguishes them. Highlighted above are the actions the prompt asks for: gpt-image-1.5 holds every face and does not perform them, wan2.7-image-pro performs them and loses the faces — the sheets never changed, yet the same white-coated woman is boxed again and again.
标注框由三名标注员各自独立画出,框色用于区分标注员。高亮处是提示词要求的动作:gpt-image-1.5 保住人脸却未做出动作,wan2.7-image-pro 做出动作却丢失身份——设定卡自始至终没变,同一位白大褂女性仍在四格中被反复框出。
We would like to express our sincere thanks to Yunxin Li, Baotian Hu and Min Zhang (Shenzhen Loop Area Institute), Yong Xiang (Peking University), Fan Hong (Beijing Film Academy) and Fei Gao (Shenzhen University) for their support on this paper. We also greatly appreciate the professional advice on film and television production from Yekai Xu and Tianlun Huang (directors), Yanyi Li (screenwriter), Jiahou Huang (producer), Jiaxin Yuan (art director), Xiang Chen, Jiayu Li and Zichen Tang (cinematographers), Ningxuan Zhang (editor), and Shangheng Jiang (colourist). We further thank Huxin Peng and Liang Dong (Tencent) for their help with data procurement.
衷心感谢深圳河套学院的李云鑫老师、户保田老师、张民老师,北京大学的向勇老师,北京电影学院的洪帆老师,深圳大学的高飞老师对本文的支持。也特别感谢导演专家徐烨凯、黄天伦,编剧专家李彦仪,制片专家黄家后,美术专家袁佳欣,摄影专家陈想、李佳雨、汤子晨,剪辑专家张宁轩,调色专家蒋尚恒在影视方面提供的专业意见。感谢腾讯的彭湖鑫、董梁在数据采购上的帮助。
The paper is on arXiv as arXiv:2609.00646 (PDF). DramaChain Bench supports extensible evaluation rather than static benchmarking: the dimension set, the judge framework and a subset of the data will follow.
论文已发布于 arXiv:arXiv:2609.00646(PDF)。DramaChain Bench 支持可扩展的持续评测,而非一次性的静态基准:维度体系、评测框架与部分数据将陆续开放。
@article{shi2026dramachain,
title = {DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation},
author = {Shi, Haoyuan and Chen, Mingtao and Jiang, Shuo and Chen, Ziyan and
Sheng, Xuyi and Liu, Yiming and Zhang, Ying and Wang, Miao and
Lu, Jianxiang and Lu, Fanyang and Lu, Songyuanyi and Wu, Xiele and
Hu, Zhichao and Liu, Yuhong and Xuan, Richeng},
journal = {arXiv preprint arXiv:2609.00646},
eprint = {2609.00646},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
year = {2026}
}