{
  "status": "complete",
  "models": [
    {
      "id": "hyperflow-8",
      "name": "HyperFlow",
      "steps": 8,
      "label": "FL2VA 训练 · 作者支持 Ref2VA",
      "video_shift": 12,
      "audio_shift": 3
    },
    {
      "id": "fasth3-4",
      "name": "FastH3",
      "steps": 4,
      "label": "T2VA 权重迁移测试",
      "video_shift": 12,
      "audio_shift": 3
    },
    {
      "id": "lightx2v-ref-8",
      "name": "LightX2V Ref2VA",
      "steps": 8,
      "label": "Ref2VA 专用 v1.0",
      "video_shift": 6,
      "audio_shift": 3
    },
    {
      "id": "lightx2v-ref-4",
      "name": "LightX2V Ref2VA",
      "steps": 4,
      "label": "Ref2VA 专用 v0.1",
      "video_shift": 6,
      "audio_shift": 3
    }
  ],
  "samples": [
    {
      "id": 1,
      "source_case": 1,
      "slug": "reference-case-1",
      "title": {
        "zh": "雨夜双角色交锋 · 人物与环境参考",
        "en": "Public Ref2VA case 1"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 5.1667-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1, 00:00.000-00:01.200] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\n[Shot 1, 00:01.200-00:03.600] Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\n[Shot 1, 00:03.600-00:05.166] They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 15-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\nDuring the middle of the same continuous shot, Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\nDuring the final phase of the same continuous shot, They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\nThe character sheets provide identity and clothing, not a collage to reproduce. Show one physical version of each fighter in the environment, with their separate silhouettes legible against the billboard. The silver blade belongs only to <Subject 1>; the blue dagger belongs only to <Subject 2>. Keep their hands connected naturally to the weapon handles and show planted feet disturbing shallow puddles. During the opening four seconds, let the rain establish depth while both women watch each other without changing sides. Across the middle six seconds, preserve the contact order of approach, slash, block and counter; the sideways camera travels only far enough to keep both bodies in view. Reserve the last five seconds for the rotation, final lock and held reaction. The final two-shot keeps both referenced faces readable, with rain, breathing and small cloak movements continuing around their otherwise stable pose.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "prompt_sha256": "5d2a08b4360b9080603ddd8fb85c73c3386a5dda88600dd557e1626eaaabedb9",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_1.jpg",
          "label": "Picture 1",
          "sha256": "1665c52d9e20ace593cbc30d1e17b5cf996ad25a59c7f1890d636711da93dbd3",
          "bytes": 344933,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_1.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_1_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_2.jpg",
          "label": "Picture 2",
          "sha256": "df77cc201252e90cb828132863d4126c4d62d512f79299ec2651bcc9be8c5b1f",
          "bytes": 396459,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_2.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_1_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_3.jpg",
          "label": "Picture 3",
          "sha256": "82cdb7a20d6ec0d7c6c396cc8e99d3f5fe49a2f67f7c3e5e75d45102f4a91062",
          "bytes": 390923,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_3.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_1_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7301,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ]
    },
    {
      "id": 2,
      "source_case": 2,
      "slug": "reference-case-2",
      "title": {
        "zh": "维多利亚书房对白 · 双角色身份",
        "en": "Public Ref2VA case 2"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 5.1667-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\n[Shot 1, 00:01.300-00:03.200] <Subject 2> leans forward just enough to make the leather chair creak, studies the detective, and asks with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\n[Shot 1, 00:03.200-00:05.166] <Subject 1> meets his eyeline and replies in a low, dry voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 15-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\nDuring the middle of the same continuous shot, <Subject 2> (S1) leans forward just enough to make the leather chair creak, studies the detective, and asks in a warm, mid-pitched English voice with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\nDuring the final phase of the same continuous shot, <Subject 1> (S2) meets his eyeline and replies in a low, dry English voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\nUse the portrait and turnaround sheets to preserve each man's identity, not to reproduce a reference-board layout. The seated doctor's sturdy shoulders, moustache and cane distinguish him from the standing detective's narrow profile and pipe. Keep their hands and props separate and their eyelines directed toward each other. The first four seconds establish their positions and the room's depth: chair legs sit firmly on the rug, the desk remains behind them, and warm reflections move subtly across the wood paneling. Allow roughly five seconds for the doctor's complete question and the detective's silent consideration, then leave the remaining six seconds for the complete reply and both reactions. Do not add another line. The doctor listens with his lips closed during the reply; the detective keeps his lips closed during the question. Small breaths, a natural blink and a restrained finger adjustment prevent a frozen tableau. Preserve the fixed viewpoint through the final pause, keeping the fire visible and the cool rainy windows distinct behind the two warm-lit faces.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "593dde5ba7be9ded40395af7de00ed01fd5cc7e1ac778276bac5c61c09f7f278",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_1.jpg",
          "label": "Picture 1",
          "sha256": "3808f06a4b21910dcc04a4f62f6fe95714d0c9773cd9f81fd67695d4eb924373",
          "bytes": 383601,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_1.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_2_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_2.jpg",
          "label": "Picture 2",
          "sha256": "b62cf2f93c256975c80fd58869216a819749a8f3aab4bc03b396b6147f0848eb",
          "bytes": 278692,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_2.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_2_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_3.jpg",
          "label": "Picture 3",
          "sha256": "d1aac91407587e4457ad191e3867e5f24dd878870ca4bbafc40ca24c605c86df",
          "bytes": 194821,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_3.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_2_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7302,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ]
    },
    {
      "id": 3,
      "source_case": 4,
      "slug": "reference-case-4",
      "title": {
        "zh": "沙漠观测者 · 单图人物与场景",
        "en": "Public Ref2VA case 4"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 5.1667-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot preserves the composition from <Picture 1>. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\n[Shot 1, 00:01.300-00:03.600] He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\n[Shot 1, 00:03.600-00:05.166] <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 15-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot keeps the referenced man, telescope, lantern and desert setting spatially coherent. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\nDuring the middle of the same continuous shot, He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\nDuring the final phase of the same continuous shot, <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\nTreat <Picture 1> as a source for the man, telescope, lantern and desert setting rather than as a required first frame. Keep a single adult man in the scene with the same cropped hair, stubble and burgundy scarf; preserve the jacket seams, brass tube proportions and dark tripod supports. His hands remain visibly attached to his arms as one hand finds the focus ring and the other steadies the instrument. The opening four seconds allow the distant light to pass across his eyes while the lantern remains steady. Over the next six seconds, his gaze follows the flash, his body turns toward the instrument, and his fingers make one deliberate focusing adjustment. During the final five seconds, he leans to the eyepiece and settles into concentrated observation. Let his scarf and jacket respond to the same gentle wind throughout. Keep the table and lantern stationary, with the red pulse changing only the light on nearby surfaces. The camera holds its viewpoint, maintaining a clear gap between the man, telescope tube and mountain skyline; no new people, additional telescopes, text or sudden camera cuts appear.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "e427f08a215a42180ea4e363c45c306819885e80729580173bb9a4f59f321da2",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_4_1.jpg",
          "label": "Picture 1",
          "sha256": "3b54a682135c5e60c3584df15754c8d54177738686a0778ca89f0a326fe50bbd",
          "bytes": 290315,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_4_1.jpg",
          "public_path": "assets/ref2va-inputs/ref2va_test_4_1.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7303,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ]
    }
  ],
  "outputs": [
    {
      "id": 1,
      "source_case": 1,
      "slug": "reference-case-1",
      "title": {
        "zh": "雨夜双角色交锋 · 人物与环境参考",
        "en": "Public Ref2VA case 1"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 5.1667-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1, 00:00.000-00:01.200] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\n[Shot 1, 00:01.200-00:03.600] Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\n[Shot 1, 00:03.600-00:05.166] They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 15-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\nDuring the middle of the same continuous shot, Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\nDuring the final phase of the same continuous shot, They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\nThe character sheets provide identity and clothing, not a collage to reproduce. Show one physical version of each fighter in the environment, with their separate silhouettes legible against the billboard. The silver blade belongs only to <Subject 1>; the blue dagger belongs only to <Subject 2>. Keep their hands connected naturally to the weapon handles and show planted feet disturbing shallow puddles. During the opening four seconds, let the rain establish depth while both women watch each other without changing sides. Across the middle six seconds, preserve the contact order of approach, slash, block and counter; the sideways camera travels only far enough to keep both bodies in view. Reserve the last five seconds for the rotation, final lock and held reaction. The final two-shot keeps both referenced faces readable, with rain, breathing and small cloak movements continuing around their otherwise stable pose.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "prompt_sha256": "5d2a08b4360b9080603ddd8fb85c73c3386a5dda88600dd557e1626eaaabedb9",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_1.jpg",
          "label": "Picture 1",
          "sha256": "1665c52d9e20ace593cbc30d1e17b5cf996ad25a59c7f1890d636711da93dbd3",
          "bytes": 344933,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_2.jpg",
          "label": "Picture 2",
          "sha256": "df77cc201252e90cb828132863d4126c4d62d512f79299ec2651bcc9be8c5b1f",
          "bytes": 396459,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_3.jpg",
          "label": "Picture 3",
          "sha256": "82cdb7a20d6ec0d7c6c396cc8e99d3f5fe49a2f67f7c3e5e75d45102f4a91062",
          "bytes": 390923,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7301,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "fasth3-4",
      "adapter_sha256": "4ce198c83132251b7fd0de2503823aa49c53983f068318f66cb19eaefb7fcc12",
      "adapter_filename": "adapter_model.safetensors",
      "adapter_training_scope": "T2VA-trained adapter transferred to Ref2VA for comparison",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 64,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 12.860172481741756,
      "cold_load_s": 447.786847113166,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 400,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 2,
        "last_step": 3,
        "video_start": 8133,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 81,
        "sequence_length": 115989,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1813,
          "sink_blocks": 81,
          "threshold_density": 0.23412,
          "effective_density": 0.31369,
          "route_count_min": 197,
          "route_count_mean": 568.72,
          "route_count_max": 1813,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/fasth3-4/01.mp4",
      "poster": "assets/ref2va-15s/fasth3-4/01.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 5928519,
        "sha256": "1859723b571f1c72b151e6dd1ffc604d641b7e5518a5fb237e41c5f063c847a8",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/fasth3-4/01.mp4",
      "dataset_path": "videos/ref2va-15s/fasth3-4/01.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 2,
      "source_case": 2,
      "slug": "reference-case-2",
      "title": {
        "zh": "维多利亚书房对白 · 双角色身份",
        "en": "Public Ref2VA case 2"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 5.1667-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\n[Shot 1, 00:01.300-00:03.200] <Subject 2> leans forward just enough to make the leather chair creak, studies the detective, and asks with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\n[Shot 1, 00:03.200-00:05.166] <Subject 1> meets his eyeline and replies in a low, dry voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 15-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\nDuring the middle of the same continuous shot, <Subject 2> (S1) leans forward just enough to make the leather chair creak, studies the detective, and asks in a warm, mid-pitched English voice with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\nDuring the final phase of the same continuous shot, <Subject 1> (S2) meets his eyeline and replies in a low, dry English voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\nUse the portrait and turnaround sheets to preserve each man's identity, not to reproduce a reference-board layout. The seated doctor's sturdy shoulders, moustache and cane distinguish him from the standing detective's narrow profile and pipe. Keep their hands and props separate and their eyelines directed toward each other. The first four seconds establish their positions and the room's depth: chair legs sit firmly on the rug, the desk remains behind them, and warm reflections move subtly across the wood paneling. Allow roughly five seconds for the doctor's complete question and the detective's silent consideration, then leave the remaining six seconds for the complete reply and both reactions. Do not add another line. The doctor listens with his lips closed during the reply; the detective keeps his lips closed during the question. Small breaths, a natural blink and a restrained finger adjustment prevent a frozen tableau. Preserve the fixed viewpoint through the final pause, keeping the fire visible and the cool rainy windows distinct behind the two warm-lit faces.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "593dde5ba7be9ded40395af7de00ed01fd5cc7e1ac778276bac5c61c09f7f278",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_1.jpg",
          "label": "Picture 1",
          "sha256": "3808f06a4b21910dcc04a4f62f6fe95714d0c9773cd9f81fd67695d4eb924373",
          "bytes": 383601,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_2.jpg",
          "label": "Picture 2",
          "sha256": "b62cf2f93c256975c80fd58869216a819749a8f3aab4bc03b396b6147f0848eb",
          "bytes": 278692,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_3.jpg",
          "label": "Picture 3",
          "sha256": "d1aac91407587e4457ad191e3867e5f24dd878870ca4bbafc40ca24c605c86df",
          "bytes": 194821,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7302,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "fasth3-4",
      "adapter_sha256": "4ce198c83132251b7fd0de2503823aa49c53983f068318f66cb19eaefb7fcc12",
      "adapter_filename": "adapter_model.safetensors",
      "adapter_training_scope": "T2VA-trained adapter transferred to Ref2VA for comparison",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 64,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 14.802989835850894,
      "cold_load_s": 447.786847113166,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 600,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 3,
        "last_step": 3,
        "video_start": 8274,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 83,
        "sequence_length": 116130,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1815,
          "sink_blocks": 83,
          "threshold_density": 0.23333,
          "effective_density": 0.31483,
          "route_count_min": 184,
          "route_count_mean": 571.42,
          "route_count_max": 1815,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/fasth3-4/02.mp4",
      "poster": "assets/ref2va-15s/fasth3-4/02.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 3351191,
        "sha256": "da74ef69309ce6c376dd73b4c94026c15cee82417dcf479ad3e092fc764c85cd",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/fasth3-4/02.mp4",
      "dataset_path": "videos/ref2va-15s/fasth3-4/02.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 3,
      "source_case": 4,
      "slug": "reference-case-4",
      "title": {
        "zh": "沙漠观测者 · 单图人物与场景",
        "en": "Public Ref2VA case 4"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 5.1667-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot preserves the composition from <Picture 1>. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\n[Shot 1, 00:01.300-00:03.600] He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\n[Shot 1, 00:03.600-00:05.166] <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 15-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot keeps the referenced man, telescope, lantern and desert setting spatially coherent. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\nDuring the middle of the same continuous shot, He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\nDuring the final phase of the same continuous shot, <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\nTreat <Picture 1> as a source for the man, telescope, lantern and desert setting rather than as a required first frame. Keep a single adult man in the scene with the same cropped hair, stubble and burgundy scarf; preserve the jacket seams, brass tube proportions and dark tripod supports. His hands remain visibly attached to his arms as one hand finds the focus ring and the other steadies the instrument. The opening four seconds allow the distant light to pass across his eyes while the lantern remains steady. Over the next six seconds, his gaze follows the flash, his body turns toward the instrument, and his fingers make one deliberate focusing adjustment. During the final five seconds, he leans to the eyepiece and settles into concentrated observation. Let his scarf and jacket respond to the same gentle wind throughout. Keep the table and lantern stationary, with the red pulse changing only the light on nearby surfaces. The camera holds its viewpoint, maintaining a clear gap between the man, telescope tube and mountain skyline; no new people, additional telescopes, text or sudden camera cuts appear.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "e427f08a215a42180ea4e363c45c306819885e80729580173bb9a4f59f321da2",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_4_1.jpg",
          "label": "Picture 1",
          "sha256": "3b54a682135c5e60c3584df15754c8d54177738686a0778ca89f0a326fe50bbd",
          "bytes": 290315,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_4_1.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7303,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "fasth3-4",
      "adapter_sha256": "4ce198c83132251b7fd0de2503823aa49c53983f068318f66cb19eaefb7fcc12",
      "adapter_filename": "adapter_model.safetensors",
      "adapter_training_scope": "T2VA-trained adapter transferred to Ref2VA for comparison",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 64,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 27.752917824778706,
      "cold_load_s": 447.786847113166,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 800,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 4,
        "last_step": 3,
        "video_start": 4009,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 49,
        "sequence_length": 111865,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1748,
          "sink_blocks": 49,
          "threshold_density": 0.22221,
          "effective_density": 0.27221,
          "route_count_min": 129,
          "route_count_mean": 475.83,
          "route_count_max": 1748,
          "metadata_capacity": 1792
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/fasth3-4/03.mp4",
      "poster": "assets/ref2va-15s/fasth3-4/03.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 3411683,
        "sha256": "c6e23a088e4ec4fb84a748859b42e250595cea0df5a72929d6a88d8700cebb0a",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/fasth3-4/03.mp4",
      "dataset_path": "videos/ref2va-15s/fasth3-4/03.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 1,
      "source_case": 1,
      "slug": "reference-case-1",
      "title": {
        "zh": "雨夜双角色交锋 · 人物与环境参考",
        "en": "Public Ref2VA case 1"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 5.1667-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1, 00:00.000-00:01.200] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\n[Shot 1, 00:01.200-00:03.600] Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\n[Shot 1, 00:03.600-00:05.166] They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 15-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\nDuring the middle of the same continuous shot, Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\nDuring the final phase of the same continuous shot, They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\nThe character sheets provide identity and clothing, not a collage to reproduce. Show one physical version of each fighter in the environment, with their separate silhouettes legible against the billboard. The silver blade belongs only to <Subject 1>; the blue dagger belongs only to <Subject 2>. Keep their hands connected naturally to the weapon handles and show planted feet disturbing shallow puddles. During the opening four seconds, let the rain establish depth while both women watch each other without changing sides. Across the middle six seconds, preserve the contact order of approach, slash, block and counter; the sideways camera travels only far enough to keep both bodies in view. Reserve the last five seconds for the rotation, final lock and held reaction. The final two-shot keeps both referenced faces readable, with rain, breathing and small cloak movements continuing around their otherwise stable pose.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "prompt_sha256": "5d2a08b4360b9080603ddd8fb85c73c3386a5dda88600dd557e1626eaaabedb9",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_1.jpg",
          "label": "Picture 1",
          "sha256": "1665c52d9e20ace593cbc30d1e17b5cf996ad25a59c7f1890d636711da93dbd3",
          "bytes": 344933,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_2.jpg",
          "label": "Picture 2",
          "sha256": "df77cc201252e90cb828132863d4126c4d62d512f79299ec2651bcc9be8c5b1f",
          "bytes": 396459,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_3.jpg",
          "label": "Picture 3",
          "sha256": "82cdb7a20d6ec0d7c6c396cc8e99d3f5fe49a2f67f7c3e5e75d45102f4a91062",
          "bytes": 390923,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7301,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "hyperflow-8",
      "adapter_sha256": "4d7dec1363ebcb9fd63117621b65f8bd19fecacf7ba41f38dd098be363d3972d",
      "adapter_filename": "minimax_h3_hyperflow_8step_v1.0.safetensors",
      "adapter_training_scope": "FL2VA-trained, author-supported Ref2VA reuse",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 256.0,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 28.191384123056196,
      "cold_load_s": 428.3224484589882,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 300,
      "sparse_stats": {
        "sparse_calls": 600,
        "dense_calls": 200,
        "sparse_calls_measured_request": 300,
        "sparse_fraction": 0.75,
        "requests": 2,
        "last_step": 7,
        "video_start": 8133,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 81,
        "sequence_length": 115989,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 2,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1813,
          "sink_blocks": 81,
          "threshold_density": 0.23208,
          "effective_density": 0.31159,
          "route_count_min": 173,
          "route_count_mean": 564.92,
          "route_count_max": 1813,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {
          "warmup_step": 200
        }
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/hyperflow-8/01.mp4",
      "poster": "assets/ref2va-15s/hyperflow-8/01.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 5283618,
        "sha256": "31c33092ce3dfe72719623309f15bafdc9511dd66f7810ef67f1ce512f9a5453",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/hyperflow-8/01.mp4",
      "dataset_path": "videos/ref2va-15s/hyperflow-8/01.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 2,
      "source_case": 2,
      "slug": "reference-case-2",
      "title": {
        "zh": "维多利亚书房对白 · 双角色身份",
        "en": "Public Ref2VA case 2"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 5.1667-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\n[Shot 1, 00:01.300-00:03.200] <Subject 2> leans forward just enough to make the leather chair creak, studies the detective, and asks with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\n[Shot 1, 00:03.200-00:05.166] <Subject 1> meets his eyeline and replies in a low, dry voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 15-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\nDuring the middle of the same continuous shot, <Subject 2> (S1) leans forward just enough to make the leather chair creak, studies the detective, and asks in a warm, mid-pitched English voice with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\nDuring the final phase of the same continuous shot, <Subject 1> (S2) meets his eyeline and replies in a low, dry English voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\nUse the portrait and turnaround sheets to preserve each man's identity, not to reproduce a reference-board layout. The seated doctor's sturdy shoulders, moustache and cane distinguish him from the standing detective's narrow profile and pipe. Keep their hands and props separate and their eyelines directed toward each other. The first four seconds establish their positions and the room's depth: chair legs sit firmly on the rug, the desk remains behind them, and warm reflections move subtly across the wood paneling. Allow roughly five seconds for the doctor's complete question and the detective's silent consideration, then leave the remaining six seconds for the complete reply and both reactions. Do not add another line. The doctor listens with his lips closed during the reply; the detective keeps his lips closed during the question. Small breaths, a natural blink and a restrained finger adjustment prevent a frozen tableau. Preserve the fixed viewpoint through the final pause, keeping the fire visible and the cool rainy windows distinct behind the two warm-lit faces.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "593dde5ba7be9ded40395af7de00ed01fd5cc7e1ac778276bac5c61c09f7f278",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_1.jpg",
          "label": "Picture 1",
          "sha256": "3808f06a4b21910dcc04a4f62f6fe95714d0c9773cd9f81fd67695d4eb924373",
          "bytes": 383601,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_2.jpg",
          "label": "Picture 2",
          "sha256": "b62cf2f93c256975c80fd58869216a819749a8f3aab4bc03b396b6147f0848eb",
          "bytes": 278692,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_3.jpg",
          "label": "Picture 3",
          "sha256": "d1aac91407587e4457ad191e3867e5f24dd878870ca4bbafc40ca24c605c86df",
          "bytes": 194821,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7302,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "hyperflow-8",
      "adapter_sha256": "4d7dec1363ebcb9fd63117621b65f8bd19fecacf7ba41f38dd098be363d3972d",
      "adapter_filename": "minimax_h3_hyperflow_8step_v1.0.safetensors",
      "adapter_training_scope": "FL2VA-trained, author-supported Ref2VA reuse",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 256.0,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 28.74764114804566,
      "cold_load_s": 428.3224484589882,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 300,
      "sparse_stats": {
        "sparse_calls": 900,
        "dense_calls": 300,
        "sparse_calls_measured_request": 300,
        "sparse_fraction": 0.75,
        "requests": 3,
        "last_step": 7,
        "video_start": 8274,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 83,
        "sequence_length": 116130,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 2,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1815,
          "sink_blocks": 83,
          "threshold_density": 0.23108,
          "effective_density": 0.31245,
          "route_count_min": 141,
          "route_count_mean": 567.1,
          "route_count_max": 1815,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {
          "warmup_step": 300
        }
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/hyperflow-8/02.mp4",
      "poster": "assets/ref2va-15s/hyperflow-8/02.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 2224091,
        "sha256": "cf8a4d20c3aa741577d2c194417fc0a26dea6a30edec9213763c44841f963896",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/hyperflow-8/02.mp4",
      "dataset_path": "videos/ref2va-15s/hyperflow-8/02.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 3,
      "source_case": 4,
      "slug": "reference-case-4",
      "title": {
        "zh": "沙漠观测者 · 单图人物与场景",
        "en": "Public Ref2VA case 4"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 5.1667-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot preserves the composition from <Picture 1>. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\n[Shot 1, 00:01.300-00:03.600] He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\n[Shot 1, 00:03.600-00:05.166] <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 15-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot keeps the referenced man, telescope, lantern and desert setting spatially coherent. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\nDuring the middle of the same continuous shot, He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\nDuring the final phase of the same continuous shot, <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\nTreat <Picture 1> as a source for the man, telescope, lantern and desert setting rather than as a required first frame. Keep a single adult man in the scene with the same cropped hair, stubble and burgundy scarf; preserve the jacket seams, brass tube proportions and dark tripod supports. His hands remain visibly attached to his arms as one hand finds the focus ring and the other steadies the instrument. The opening four seconds allow the distant light to pass across his eyes while the lantern remains steady. Over the next six seconds, his gaze follows the flash, his body turns toward the instrument, and his fingers make one deliberate focusing adjustment. During the final five seconds, he leans to the eyepiece and settles into concentrated observation. Let his scarf and jacket respond to the same gentle wind throughout. Keep the table and lantern stationary, with the red pulse changing only the light on nearby surfaces. The camera holds its viewpoint, maintaining a clear gap between the man, telescope tube and mountain skyline; no new people, additional telescopes, text or sudden camera cuts appear.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "e427f08a215a42180ea4e363c45c306819885e80729580173bb9a4f59f321da2",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_4_1.jpg",
          "label": "Picture 1",
          "sha256": "3b54a682135c5e60c3584df15754c8d54177738686a0778ca89f0a326fe50bbd",
          "bytes": 290315,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_4_1.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7303,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "hyperflow-8",
      "adapter_sha256": "4d7dec1363ebcb9fd63117621b65f8bd19fecacf7ba41f38dd098be363d3972d",
      "adapter_filename": "minimax_h3_hyperflow_8step_v1.0.safetensors",
      "adapter_training_scope": "FL2VA-trained, author-supported Ref2VA reuse",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 12.0,
      "audio_shift": 3.0,
      "adapter_alpha": 256.0,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 27.16704279894475,
      "cold_load_s": 428.3224484589882,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 300,
      "sparse_stats": {
        "sparse_calls": 1200,
        "dense_calls": 400,
        "sparse_calls_measured_request": 300,
        "sparse_fraction": 0.75,
        "requests": 4,
        "last_step": 7,
        "video_start": 4009,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 49,
        "sequence_length": 111865,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 2,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1748,
          "sink_blocks": 49,
          "threshold_density": 0.21892,
          "effective_density": 0.26873,
          "route_count_min": 90,
          "route_count_mean": 469.74,
          "route_count_max": 1748,
          "metadata_capacity": 1792
        },
        "gate": null,
        "declined": {
          "warmup_step": 400
        }
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/hyperflow-8/03.mp4",
      "poster": "assets/ref2va-15s/hyperflow-8/03.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 2043297,
        "sha256": "d160ccf140f4f167da93c682d4085138fb236a02cbb50eab8602223e4db4115c",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/hyperflow-8/03.mp4",
      "dataset_path": "videos/ref2va-15s/hyperflow-8/03.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 1,
      "source_case": 1,
      "slug": "reference-case-1",
      "title": {
        "zh": "雨夜双角色交锋 · 人物与环境参考",
        "en": "Public Ref2VA case 1"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 5.1667-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1, 00:00.000-00:01.200] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\n[Shot 1, 00:01.200-00:03.600] Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\n[Shot 1, 00:03.600-00:05.166] They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 15-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\nDuring the middle of the same continuous shot, Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\nDuring the final phase of the same continuous shot, They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\nThe character sheets provide identity and clothing, not a collage to reproduce. Show one physical version of each fighter in the environment, with their separate silhouettes legible against the billboard. The silver blade belongs only to <Subject 1>; the blue dagger belongs only to <Subject 2>. Keep their hands connected naturally to the weapon handles and show planted feet disturbing shallow puddles. During the opening four seconds, let the rain establish depth while both women watch each other without changing sides. Across the middle six seconds, preserve the contact order of approach, slash, block and counter; the sideways camera travels only far enough to keep both bodies in view. Reserve the last five seconds for the rotation, final lock and held reaction. The final two-shot keeps both referenced faces readable, with rain, breathing and small cloak movements continuing around their otherwise stable pose.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "prompt_sha256": "5d2a08b4360b9080603ddd8fb85c73c3386a5dda88600dd557e1626eaaabedb9",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_1.jpg",
          "label": "Picture 1",
          "sha256": "1665c52d9e20ace593cbc30d1e17b5cf996ad25a59c7f1890d636711da93dbd3",
          "bytes": 344933,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_2.jpg",
          "label": "Picture 2",
          "sha256": "df77cc201252e90cb828132863d4126c4d62d512f79299ec2651bcc9be8c5b1f",
          "bytes": 396459,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_3.jpg",
          "label": "Picture 3",
          "sha256": "82cdb7a20d6ec0d7c6c396cc8e99d3f5fe49a2f67f7c3e5e75d45102f4a91062",
          "bytes": 390923,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7301,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-4",
      "adapter_sha256": "9e642fc8749c74f8da5e2382877ab5c7aa37b9a73b7fd0d6d457bd1b3cb1ae99",
      "adapter_filename": "minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 13.126003897748888,
      "cold_load_s": 412.2830279250629,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 400,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 2,
        "last_step": 3,
        "video_start": 8133,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 81,
        "sequence_length": 115989,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1813,
          "sink_blocks": 81,
          "threshold_density": 0.23451,
          "effective_density": 0.31405,
          "route_count_min": 193,
          "route_count_mean": 569.38,
          "route_count_max": 1813,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/lightx2v-ref-4/01.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-4/01.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 5326358,
        "sha256": "34be4059e0925b3242a8706ab0078d14d1044f3b099ef96043aaa45e76a2c40b",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-4/01.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-4/01.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 2,
      "source_case": 2,
      "slug": "reference-case-2",
      "title": {
        "zh": "维多利亚书房对白 · 双角色身份",
        "en": "Public Ref2VA case 2"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 5.1667-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\n[Shot 1, 00:01.300-00:03.200] <Subject 2> leans forward just enough to make the leather chair creak, studies the detective, and asks with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\n[Shot 1, 00:03.200-00:05.166] <Subject 1> meets his eyeline and replies in a low, dry voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 15-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\nDuring the middle of the same continuous shot, <Subject 2> (S1) leans forward just enough to make the leather chair creak, studies the detective, and asks in a warm, mid-pitched English voice with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\nDuring the final phase of the same continuous shot, <Subject 1> (S2) meets his eyeline and replies in a low, dry English voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\nUse the portrait and turnaround sheets to preserve each man's identity, not to reproduce a reference-board layout. The seated doctor's sturdy shoulders, moustache and cane distinguish him from the standing detective's narrow profile and pipe. Keep their hands and props separate and their eyelines directed toward each other. The first four seconds establish their positions and the room's depth: chair legs sit firmly on the rug, the desk remains behind them, and warm reflections move subtly across the wood paneling. Allow roughly five seconds for the doctor's complete question and the detective's silent consideration, then leave the remaining six seconds for the complete reply and both reactions. Do not add another line. The doctor listens with his lips closed during the reply; the detective keeps his lips closed during the question. Small breaths, a natural blink and a restrained finger adjustment prevent a frozen tableau. Preserve the fixed viewpoint through the final pause, keeping the fire visible and the cool rainy windows distinct behind the two warm-lit faces.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "593dde5ba7be9ded40395af7de00ed01fd5cc7e1ac778276bac5c61c09f7f278",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_1.jpg",
          "label": "Picture 1",
          "sha256": "3808f06a4b21910dcc04a4f62f6fe95714d0c9773cd9f81fd67695d4eb924373",
          "bytes": 383601,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_2.jpg",
          "label": "Picture 2",
          "sha256": "b62cf2f93c256975c80fd58869216a819749a8f3aab4bc03b396b6147f0848eb",
          "bytes": 278692,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_3.jpg",
          "label": "Picture 3",
          "sha256": "d1aac91407587e4457ad191e3867e5f24dd878870ca4bbafc40ca24c605c86df",
          "bytes": 194821,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7302,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-4",
      "adapter_sha256": "9e642fc8749c74f8da5e2382877ab5c7aa37b9a73b7fd0d6d457bd1b3cb1ae99",
      "adapter_filename": "minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 13.445603016298264,
      "cold_load_s": 412.2830279250629,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 600,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 3,
        "last_step": 3,
        "video_start": 8274,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 83,
        "sequence_length": 116130,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1815,
          "sink_blocks": 83,
          "threshold_density": 0.23385,
          "effective_density": 0.31533,
          "route_count_min": 166,
          "route_count_mean": 572.32,
          "route_count_max": 1815,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/lightx2v-ref-4/02.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-4/02.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 2270858,
        "sha256": "9abc5822e9f351e71f9dcadb903bc6d57f4dadd6260a29d694aa6f7902e59b3c",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-4/02.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-4/02.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 3,
      "source_case": 4,
      "slug": "reference-case-4",
      "title": {
        "zh": "沙漠观测者 · 单图人物与场景",
        "en": "Public Ref2VA case 4"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 5.1667-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot preserves the composition from <Picture 1>. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\n[Shot 1, 00:01.300-00:03.600] He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\n[Shot 1, 00:03.600-00:05.166] <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 15-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot keeps the referenced man, telescope, lantern and desert setting spatially coherent. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\nDuring the middle of the same continuous shot, He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\nDuring the final phase of the same continuous shot, <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\nTreat <Picture 1> as a source for the man, telescope, lantern and desert setting rather than as a required first frame. Keep a single adult man in the scene with the same cropped hair, stubble and burgundy scarf; preserve the jacket seams, brass tube proportions and dark tripod supports. His hands remain visibly attached to his arms as one hand finds the focus ring and the other steadies the instrument. The opening four seconds allow the distant light to pass across his eyes while the lantern remains steady. Over the next six seconds, his gaze follows the flash, his body turns toward the instrument, and his fingers make one deliberate focusing adjustment. During the final five seconds, he leans to the eyepiece and settles into concentrated observation. Let his scarf and jacket respond to the same gentle wind throughout. Keep the table and lantern stationary, with the red pulse changing only the light on nearby surfaces. The camera holds its viewpoint, maintaining a clear gap between the man, telescope tube and mountain skyline; no new people, additional telescopes, text or sudden camera cuts appear.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "e427f08a215a42180ea4e363c45c306819885e80729580173bb9a4f59f321da2",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_4_1.jpg",
          "label": "Picture 1",
          "sha256": "3b54a682135c5e60c3584df15754c8d54177738686a0778ca89f0a326fe50bbd",
          "bytes": 290315,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_4_1.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7303,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-4",
      "adapter_sha256": "9e642fc8749c74f8da5e2382877ab5c7aa37b9a73b7fd0d6d457bd1b3cb1ae99",
      "adapter_filename": "minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 4,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 13.040617477614433,
      "cold_load_s": 412.2830279250629,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 200,
      "sparse_stats": {
        "sparse_calls": 800,
        "dense_calls": 0,
        "sparse_calls_measured_request": 200,
        "sparse_fraction": 1.0,
        "requests": 4,
        "last_step": 3,
        "video_start": 4009,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 49,
        "sequence_length": 111865,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 200,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 200,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1748,
          "sink_blocks": 49,
          "threshold_density": 0.22208,
          "effective_density": 0.27202,
          "route_count_min": 116,
          "route_count_mean": 475.5,
          "route_count_max": 1748,
          "metadata_capacity": 1792
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/c58e13028b3ba29f63b1979e9fb1622e5e4d10d2/videos/ref2va-15s/lightx2v-ref-4/03.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-4/03.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 2167906,
        "sha256": "49bee7a8b4fdba4ef810bcb6a474c35ae4406564d48d82f602635504123d9b44",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-4/03.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-4/03.mp4",
      "media_dataset_revision": "c58e13028b3ba29f63b1979e9fb1622e5e4d10d2"
    },
    {
      "id": 1,
      "source_case": 1,
      "slug": "reference-case-1",
      "title": {
        "zh": "雨夜双角色交锋 · 人物与环境参考",
        "en": "Public Ref2VA case 1"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 5.1667-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1, 00:00.000-00:01.200] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\n[Shot 1, 00:01.200-00:03.600] Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\n[Shot 1, 00:03.600-00:05.166] They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the long-haired East Asian female hunter in <Picture 1>, preserving her face, long wet black hair, black leather combat suit, tattered black cloak, black gloves, and single silver short blade.\n<Subject 2> is the East Asian female tech hunter in <Picture 2>, preserving her face, short black hair, white fox mask with red markings, black trench coat with cyan circuit lines, black boots, and single blue energy dagger.\n<Subject 3> is the rainy cyberpunk intersection in <Picture 3>, preserving its skyscrapers, giant blue-purple holographic billboard, neon signs, wet reflective asphalt, rain, and drifting fog.\n\nsummary:\n[reference generation] In one continuous 15-second cinematic shot, <Subject 1> and <Subject 2> begin in a tense standoff inside <Subject 3>, charge across the wet street, exchange one fast blade combination, and end locked blade-to-blade in a burst of blue energy.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain the same face, long hair, black combat suit, cloak, body proportions, and one silver blade.\n<Subject 2> (appears throughout): fully_preserved - retain the same face, short hair, fox mask, cyan-lit trench coat, body proportions, and one energy dagger.\n<Subject 3> (appears throughout): fully_preserved - retain the wet intersection, billboard, neon palette, rain, fog, and reflective pavement.\n\ndetailed_description:\nPhotorealistic live-action cyberpunk action cinema, 16:9, high contrast purple-and-cyan night lighting, heavy rain, physically coherent motion, and stable identities. One continuous shot with no cuts, no duplicate people, no duplicate weapons, no costume changes, no gore, and no readable text.\n\n[Shot 1] A low wide camera frames both fighters five meters apart on the rain-soaked street. <Subject 1> crouches on frame left with her silver blade held low; <Subject 2> stands on frame right with the blue energy dagger raised. Rain streaks through the holographic light and fog crosses the reflective road.\n\nDuring the middle of the same continuous shot, Both women burst forward. The camera tracks sideways at waist height as <Subject 1>'s cloak opens behind her. She makes one upward diagonal slash; <Subject 2> sidesteps and blocks with the energy dagger. A compact blue-white spark and a ring of displaced droplets mark the impact. <Subject 2> answers with one horizontal counter, which <Subject 1> deflects with the flat of her blade.\n\nDuring the final phase of the same continuous shot, They rotate once around their joined weapons and stop in a close blade lock, faces and costumes still distinct. Cyan electricity crawls briefly across the crossed blades while rain runs from the fox mask and black cloak. The camera settles into a medium two-shot as the billboard flickers behind them and both hold the unresolved confrontation.\n\nThe character sheets provide identity and clothing, not a collage to reproduce. Show one physical version of each fighter in the environment, with their separate silhouettes legible against the billboard. The silver blade belongs only to <Subject 1>; the blue dagger belongs only to <Subject 2>. Keep their hands connected naturally to the weapon handles and show planted feet disturbing shallow puddles. During the opening four seconds, let the rain establish depth while both women watch each other without changing sides. Across the middle six seconds, preserve the contact order of approach, slash, block and counter; the sideways camera travels only far enough to keep both bodies in view. Reserve the last five seconds for the rotation, final lock and held reaction. The final two-shot keeps both referenced faces readable, with rain, breathing and small cloak movements continuing around their otherwise stable pose.\n\noverall_soundscape:\nHeavy rain, wet footfalls, two sharp blade impacts, controlled fabric movement, electrical crackle, distant traffic, and a low thunder roll. No dialogue.\n\nnon_diegetic_music:\nA dark electronic pulse rises through the charge and stops on the final blade lock.\n",
      "prompt_sha256": "5d2a08b4360b9080603ddd8fb85c73c3386a5dda88600dd557e1626eaaabedb9",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_1.jpg",
          "label": "Picture 1",
          "sha256": "1665c52d9e20ace593cbc30d1e17b5cf996ad25a59c7f1890d636711da93dbd3",
          "bytes": 344933,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_2.jpg",
          "label": "Picture 2",
          "sha256": "df77cc201252e90cb828132863d4126c4d62d512f79299ec2651bcc9be8c5b1f",
          "bytes": 396459,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_1_3.jpg",
          "label": "Picture 3",
          "sha256": "82cdb7a20d6ec0d7c6c396cc8e99d3f5fe49a2f67f7c3e5e75d45102f4a91062",
          "bytes": 390923,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_1_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7301,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-8",
      "adapter_sha256": "9bac880b1a5d7ac052171cf6cce769f0cceaaa42ffa51de4b8e41143a2bdd2d2",
      "adapter_filename": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 23.357075357809663,
      "cold_load_s": 382.9373801648617,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 400,
      "sparse_stats": {
        "sparse_calls": 800,
        "dense_calls": 0,
        "sparse_calls_measured_request": 400,
        "sparse_fraction": 1.0,
        "requests": 2,
        "last_step": 7,
        "video_start": 8133,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 81,
        "sequence_length": 115989,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1813,
          "sink_blocks": 81,
          "threshold_density": 0.23447,
          "effective_density": 0.31402,
          "route_count_min": 193,
          "route_count_mean": 569.31,
          "route_count_max": 1813,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/87eca1622a36826aaad99a68a8556cf89e50ecb2/videos/ref2va-15s/lightx2v-ref-8/01.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-8/01.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 5383420,
        "sha256": "b313b22fb17717810e939b57b9a0fc5df6aba9786709997156965ea30bd7ec9f",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-8/01.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-8/01.mp4",
      "media_dataset_revision": "87eca1622a36826aaad99a68a8556cf89e50ecb2"
    },
    {
      "id": 2,
      "source_case": 2,
      "slug": "reference-case-2",
      "title": {
        "zh": "维多利亚书房对白 · 双角色身份",
        "en": "Public Ref2VA case 2"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 5.1667-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\n[Shot 1, 00:01.300-00:03.200] <Subject 2> leans forward just enough to make the leather chair creak, studies the detective, and asks with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\n[Shot 1, 00:03.200-00:05.166] <Subject 1> meets his eyeline and replies in a low, dry voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is the fictional Victorian detective in <Picture 3>, preserving his angular clean-shaven face, combed-back dark hair, long dark overcoat, brown waistcoat, cream cravat, watch chain, and briar pipe.\n<Subject 2> is the fictional Victorian doctor in <Picture 2>, preserving his sturdy build, cropped brown hair, heavy moustache, brown tweed suit, dark tie, and wooden cane.\n<Subject 3> is the Baker Street study in <Picture 1>, preserving the wood paneling, lit fireplace, bookshelves, rain-streaked windows, leather armchair, desk, chandelier, rug, and warm-firelight-versus-cool-window-light composition.\n\nsummary:\n[reference generation] In one continuous 15-second locked-off Victorian dialogue shot, <Subject 2> asks whether <Subject 1> could have become a criminal; <Subject 1> answers with dry confidence while both remain naturally alive in <Subject 3>.\n\nretention_analysis:\n<Subject 1> (appears throughout): fully_preserved - retain his face, hair, clothing, pipe, slim proportions, and fireplace-side placement.\n<Subject 2> (appears throughout): fully_preserved - retain his face, moustache, tweed suit, cane, sturdy proportions, and armchair placement.\n<Subject 3> (appears throughout): fully_preserved - retain the room layout, fireplace, shelves, windows, desk, furniture, rain, and mixed warm-cool lighting.\n\ndetailed_description:\nPhotorealistic live-action Victorian feature-film style, restrained 35 mm texture, natural skin, subtle performance, and consistent eyelines. One continuous medium-wide two-shot from a fully static locked-off camera. No cuts, camera movement, identity swapping, wardrobe changes, extra people, readable text, or exaggerated gestures.\n\n[Shot 1] <Subject 1> stands beside the fireplace on frame left, his pipe held near his chest, while <Subject 2> sits in the leather armchair on frame right with one hand resting on his cane. Wind-driven rain traces the cool windows behind them. A coal shifts in the grate, briefly sharpening warm firelight across <Subject 1>'s angular profile and <Subject 2>'s moustached face.\n\nDuring the middle of the same continuous shot, <Subject 2> (S1) leans forward just enough to make the leather chair creak, studies the detective, and asks in a warm, mid-pitched English voice with familiar curiosity, <d>[English] What if you'd chosen crime?</d> His fingers tighten once around the cane handle. <Subject 1> remains still for a beat, draws quietly from the pipe, and turns only his eyes toward the doctor as smoke curls through the warm light.\n\nDuring the final phase of the same continuous shot, <Subject 1> (S2) meets his eyeline and replies in a low, dry English voice, <d>[English] I'd have been remarkably successful.</d> <Subject 2>'s amused expression tightens into thoughtful concern. <Subject 1> allows the faintest private smile as the fire gives one crisp crackle and a distant lightning reflection passes across the rain-streaked glass. Both settle into the charged silence without changing position.\n\nUse the portrait and turnaround sheets to preserve each man's identity, not to reproduce a reference-board layout. The seated doctor's sturdy shoulders, moustache and cane distinguish him from the standing detective's narrow profile and pipe. Keep their hands and props separate and their eyelines directed toward each other. The first four seconds establish their positions and the room's depth: chair legs sit firmly on the rug, the desk remains behind them, and warm reflections move subtly across the wood paneling. Allow roughly five seconds for the doctor's complete question and the detective's silent consideration, then leave the remaining six seconds for the complete reply and both reactions. Do not add another line. The doctor listens with his lips closed during the reply; the detective keeps his lips closed during the question. Small breaths, a natural blink and a restrained finger adjustment prevent a frozen tableau. Preserve the fixed viewpoint through the final pause, keeping the fire visible and the cool rainy windows distinct behind the two warm-lit faces.\n\noverall_soundscape:\nSteady rain on glass, close coal-fire crackle, faint room tone, one subtle leather creak, and distant carriage wheels. Dialogue stays clear and intimate.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "593dde5ba7be9ded40395af7de00ed01fd5cc7e1ac778276bac5c61c09f7f278",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_1.jpg",
          "label": "Picture 1",
          "sha256": "3808f06a4b21910dcc04a4f62f6fe95714d0c9773cd9f81fd67695d4eb924373",
          "bytes": 383601,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_1.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_2.jpg",
          "label": "Picture 2",
          "sha256": "b62cf2f93c256975c80fd58869216a819749a8f3aab4bc03b396b6147f0848eb",
          "bytes": 278692,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_2.jpg"
        },
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_2_3.jpg",
          "label": "Picture 3",
          "sha256": "d1aac91407587e4457ad191e3867e5f24dd878870ca4bbafc40ca24c605c86df",
          "bytes": 194821,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_2_3.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7302,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-8",
      "adapter_sha256": "9bac880b1a5d7ac052171cf6cce769f0cceaaa42ffa51de4b8e41143a2bdd2d2",
      "adapter_filename": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 23.864395305048674,
      "cold_load_s": 382.9373801648617,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 400,
      "sparse_stats": {
        "sparse_calls": 1200,
        "dense_calls": 0,
        "sparse_calls_measured_request": 400,
        "sparse_fraction": 1.0,
        "requests": 3,
        "last_step": 7,
        "video_start": 8274,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 83,
        "sequence_length": 116130,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1815,
          "sink_blocks": 83,
          "threshold_density": 0.23381,
          "effective_density": 0.31529,
          "route_count_min": 166,
          "route_count_mean": 572.25,
          "route_count_max": 1815,
          "metadata_capacity": 1856
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/87eca1622a36826aaad99a68a8556cf89e50ecb2/videos/ref2va-15s/lightx2v-ref-8/02.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-8/02.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 2140176,
        "sha256": "0fcbf81a6f495c17c174cf98ccea176b96e71171a3486e6824c564440b3da440",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-8/02.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-8/02.mp4",
      "media_dataset_revision": "87eca1622a36826aaad99a68a8556cf89e50ecb2"
    },
    {
      "id": 3,
      "source_case": 4,
      "slug": "reference-case-4",
      "title": {
        "zh": "沙漠观测者 · 单图人物与场景",
        "en": "Public Ref2VA case 4"
      },
      "source": {
        "name": "LightX2V public Ref2VA testset",
        "url": "https://github.com/ModelTC/Minimax-H3-Turbo/blob/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/prompts_ref2va_test.json"
      },
      "source_prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 5.1667-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot preserves the composition from <Picture 1>. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1, 00:00.000-00:01.300] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\n[Shot 1, 00:01.300-00:03.600] He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\n[Shot 1, 00:03.600-00:05.166] <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "source_duration": 5.1666666667,
      "prompt": "subject_definitions:\n<Subject 1> is a tan-skinned adult man with close-cropped dark hair, light stubble, a sand-colored field jacket, and a burgundy scarf, as shown in <Picture 1>.\n<Subject 2> is a vintage brass telescope on a dark tripod, as shown in <Picture 1>.\n<Subject 3> is a small red-glowing metal lantern on a weathered table, as shown in <Picture 1>.\n<Subject 4> is a quiet desert observatory at dusk with soft mountain silhouettes, as shown in <Picture 1>.\n\nsummary:\n[reference generation] In one continuous 15-second shot inside <Subject 4>, <Subject 1> notices a distant flash, turns toward <Subject 2>, adjusts its focus ring once, and reaches the eyepiece as <Subject 3> pulses red.\n\nretention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - Preserve the same adult identity, facial structure, hair, stubble, jacket, scarf, and body proportions.\n<Subject 2> (appears in [Shot 1]): fully_preserved - Preserve the same brass tube, black fittings, tripod geometry, scale, and position.\n<Subject 3> (appears in [Shot 1]): fully_preserved - Preserve the same metal housing and red light.\n<Subject 4> (appears in [Shot 1]): fully_preserved - Preserve the dusk palette, mountain horizon, and sparse desert setting.\n\ndetailed_description:\nVisual style: cinematic naturalism with warm lantern light against a cool violet dusk. One locked medium shot keeps the referenced man, telescope, lantern and desert setting spatially coherent. The camera stays fixed with no pan, tilt, zoom, or dolly.\n\n[Shot 1] <Subject 1> stands beside <Subject 2> as the final dusk light outlines the brass tube and distant mountain ridge. A faint cold flash crosses his eyes from outside frame. He stops breathing for a beat and turns his gaze sharply toward the telescope while <Subject 3> holds a steady red glow on the table.\n\nDuring the middle of the same continuous shot, He reaches across the brass tube, grips the focus ring, and makes exactly one deliberate adjustment. The mechanism gives a precise metal click. Reflected violet sky slides across the brass surface, his scarf lifts in the desert wind, and <Subject 3> flickers once brighter without moving.\n\nDuring the final phase of the same continuous shot, <Subject 1> leans quickly but naturally to the eyepiece, steadies the telescope with his free hand, and peers through it with sudden concentration. The lantern throws a brief red edge across his cheek and telescope fittings. He holds the tense viewing pose as the distant flash fades behind the mountains.\n\nTreat <Picture 1> as a source for the man, telescope, lantern and desert setting rather than as a required first frame. Keep a single adult man in the scene with the same cropped hair, stubble and burgundy scarf; preserve the jacket seams, brass tube proportions and dark tripod supports. His hands remain visibly attached to his arms as one hand finds the focus ring and the other steadies the instrument. The opening four seconds allow the distant light to pass across his eyes while the lantern remains steady. Over the next six seconds, his gaze follows the flash, his body turns toward the instrument, and his fingers make one deliberate focusing adjustment. During the final five seconds, he leans to the eyepiece and settles into concentrated observation. Let his scarf and jacket respond to the same gentle wind throughout. Keep the table and lantern stationary, with the red pulse changing only the light on nearby surfaces. The camera holds its viewpoint, maintaining a clear gap between the man, telescope tube and mountain skyline; no new people, additional telescopes, text or sudden camera cuts appear.\n\noverall_soundscape:\nA faint desert breeze, a soft metal adjustment click, fabric rustle, and a quiet lantern flame.\n\nnon_diegetic_music:\nN/A\n",
      "prompt_sha256": "e427f08a215a42180ea4e363c45c306819885e80729580173bb9a4f59f321da2",
      "references": [
        {
          "type": "image",
          "path": "ref-inputs/ref2va_testset/ref2va_test_4_1.jpg",
          "label": "Picture 1",
          "sha256": "3b54a682135c5e60c3584df15754c8d54177738686a0778ca89f0a326fe50bbd",
          "bytes": 290315,
          "source_url": "https://raw.githubusercontent.com/ModelTC/Minimax-H3-Turbo/02e26d591f7a04d5d1a074c9566d5dd4f22f6225/examples/ref2va_testset/ref2va_test_4_1.jpg"
        }
      ],
      "duration": 15,
      "frames": 362,
      "seed": 7303,
      "adaptations": [
        "保留上游六字段提示词与原始参考图片；时间安排由约 5.17 秒扩展为 15 秒。",
        "将上游重复的 Shot 1 时间段整理为一个连续镜头；明确对白说话人；原文台词未改。",
        "新增动作衔接与参考保持说明属于本次整理；本组仅测试图片参考，声音由模型生成。"
      ],
      "adapter": "lightx2v-ref-8",
      "adapter_sha256": "9bac880b1a5d7ac052171cf6cce769f0cceaaa42ffa51de4b8e41143a2bdd2d2",
      "adapter_filename": "minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors",
      "adapter_training_scope": "Native Ref2VA distilled adapter",
      "task": "ref2va",
      "nfe": 8,
      "cluster": "hsg",
      "gpus": 4,
      "nodes": 1,
      "gpu": "NVIDIA GB200",
      "compute_quant": "none",
      "attention_backend": "sol_bsa",
      "width": 1344,
      "height": 768,
      "fps": 24,
      "video_shift": 6.0,
      "audio_shift": 3.0,
      "adapter_alpha": 8,
      "reference_image_resize_mode": "match",
      "reference_conditioning_schedule": 1,
      "pipeline_s": 22.327419661916792,
      "cold_load_s": 382.9373801648617,
      "base_model_revision": "83db0c0efe6ef9824e0e194be110346c0a9542ed",
      "observed_sparse_calls": 400,
      "sparse_stats": {
        "sparse_calls": 1600,
        "dense_calls": 0,
        "sparse_calls_measured_request": 400,
        "sparse_fraction": 1.0,
        "requests": 4,
        "last_step": 7,
        "video_start": 4009,
        "sink_mode": "text_audio",
        "sink_start": 0,
        "sink_tokens": 0,
        "sink_blocks": 49,
        "sequence_length": 111865,
        "tau": 1.0,
        "thresh_type": "diag",
        "backend": "sol_bsa",
        "qkv_wire_dtype": "int8_qkv",
        "qkv_wire_policy": "int8_qkv_all",
        "qkv_int8_calls_measured_request": 400,
        "qkv_bf16_calls_measured_request": 0,
        "qkv_wire_ratio": 0.520833,
        "output_wire_dtype": "fp8",
        "output_int8_calls_measured_request": 0,
        "output_fp8_calls_measured_request": 400,
        "output_bf16_calls_measured_request": 0,
        "output_wire_ratio": 0.5,
        "attention_comm_ratio": 0.515625,
        "packed_input": true,
        "compile_bucket_size": 4096,
        "dense_steps": 0,
        "dense_layers": 0,
        "route_density": {
          "blocks": 1748,
          "sink_blocks": 49,
          "threshold_density": 0.22207,
          "effective_density": 0.27201,
          "route_count_min": 115,
          "route_count_mean": 475.47,
          "route_count_max": 1748,
          "metadata_capacity": 1792
        },
        "gate": null,
        "declined": {}
      },
      "timing": "Single prompt request including reference/text encoding, denoising and video/audio decoding; MP4 encoding excluded",
      "video": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media/resolve/87eca1622a36826aaad99a68a8556cf89e50ecb2/videos/ref2va-15s/lightx2v-ref-8/03.mp4",
      "poster": "assets/ref2va-15s/lightx2v-ref-8/03.jpg",
      "media_validation": {
        "video_codec": "h264",
        "audio_codec": "aac",
        "frames": 362,
        "width": 1344,
        "height": 768,
        "fps": 24,
        "sample_rate": 32000,
        "channels": 2,
        "duration_s": 15.166666,
        "bytes": 1847958,
        "sha256": "4db3586b516a00f22648401f3baa5dae80598ed1307214550a0c94278fe20f2b",
        "verification": "ffprobe frame count and complete audio/video decode"
      },
      "local_video": "assets/ref2va-15s/lightx2v-ref-8/03.mp4",
      "dataset_path": "videos/ref2va-15s/lightx2v-ref-8/03.mp4",
      "media_dataset_revision": "87eca1622a36826aaad99a68a8556cf89e50ecb2"
    }
  ],
  "source_revision": "02e26d591f7a04d5d1a074c9566d5dd4f22f6225",
  "updated_at": "2026-09-13T10:19:10.477601+00:00",
  "profile": {
    "task": "ref2va",
    "duration_s": 15,
    "frames": 362,
    "gpus": 4,
    "cluster": "hsg",
    "compute_quant": "none",
    "reference_types_tested": [
      "image"
    ]
  },
  "media_repository": {
    "repo_id": "Lawrence-cj/sol-h3-hyperflow-20260913-media",
    "latest_commit": "87eca1622a36826aaad99a68a8556cf89e50ecb2",
    "url": "https://huggingface.co/datasets/Lawrence-cj/sol-h3-hyperflow-20260913-media"
  }
}
