{ jsonpromptstudio }

Street Interviews · "Public Freakout Incoming!"

Downtown Chaos

Ask strangers wild questions and capture unfiltered reactions. Replace the {{highlighted}} variables below with your own twist — suggestions included — then copy the JSON for your model.

Google · duration 4s, 6s, or 8s (8s required for 1080p, 4K, or reference images) · 720p, 1080p, or 4K · audio native · verified 2026-07-11 - click any value to edit
{
  "scene": {
    "location": "busy downtown street with heavy foot traffic",
    "mood": "chaotic street energy with unexpected encounters",
    "environment": "bustling city street with ambient urban sounds"
  },
  "action": "approaching strangers, quick interviews, reactions and walk-aways",
  "camera": {
    "angle": "handheld street-level shots",
    "distance": "medium shots capturing both interviewer and subject",
    "movement": "mobile, following the action"
  },
  "lighting": {
    "type": "natural daylight with urban environment"
  },
  "style": {
    "tone": "energetic and unpredictable",
    "color_palette": "vibrant urban colors",
    "location": "{{location}}",
    "interview_question": "{{interview_question}}",
    "sample_answers": "{{sample_answers}}",
    "tone_style": "{{tone_style}}"
  },
  "audio": {
    "music": "upbeat street energy background"
  }
}
  • Veo 3.1 follows structured prompts unusually well: lock camera, lighting and audio as separate JSON fields.
  • Native audio: describe dialogue, ambient sound and music directly in the audio field.
Kuaishou · duration expanded limits; exact caps vary by mode · mode-dependent · audio native audio-visual output · verified 2026-07-11 - click any value to edit
{
  "action": "approaching strangers, quick interviews, reactions and walk-aways",
  "scene": {
    "location": "busy downtown street with heavy foot traffic",
    "mood": "chaotic street energy with unexpected encounters",
    "environment": "bustling city street with ambient urban sounds"
  },
  "camera": {
    "angle": "handheld street-level shots",
    "distance": "medium shots capturing both interviewer and subject",
    "movement": "mobile, following the action"
  },
  "lighting": {
    "type": "natural daylight with urban environment"
  },
  "style": {
    "tone": "energetic and unpredictable",
    "color_palette": "vibrant urban colors",
    "location": "{{location}}",
    "interview_question": "{{interview_question}}",
    "sample_answers": "{{sample_answers}}",
    "tone_style": "{{tone_style}}"
  },
  "audio": {
    "music": "upbeat street energy background"
  }
}
  • Kling prompt guidance emphasizes Subject, Movement, Scene, Camera Language and Lighting.
  • Write the action field as a shot note, not a keyword list.
ByteDance · duration 4-15s · 480p, 720p, 1080p, 4K ratios in Runway API · audio native synchronized audio on supported hosts · verified 2026-07-11 - click any value to edit
{
  "camera": {
    "angle": "handheld street-level shots",
    "distance": "medium shots capturing both interviewer and subject",
    "movement": "mobile, following the action"
  },
  "action": "approaching strangers, quick interviews, reactions and walk-aways",
  "scene": {
    "location": "busy downtown street with heavy foot traffic",
    "mood": "chaotic street energy with unexpected encounters",
    "environment": "bustling city street with ambient urban sounds"
  },
  "lighting": {
    "type": "natural daylight with urban environment"
  },
  "style": {
    "tone": "energetic and unpredictable",
    "color_palette": "vibrant urban colors",
    "location": "{{location}}",
    "interview_question": "{{interview_question}}",
    "sample_answers": "{{sample_answers}}",
    "tone_style": "{{tone_style}}"
  },
  "audio": {
    "music": "upbeat street energy background"
  }
}
  • Use a strict block order: CAMERA -> SUBJECT -> ACTION -> ENVIRONMENT -> LIGHTING -> STYLE.
  • Keep the camera block to a shot type plus one movement; stacking moves degrades output.
Alibaba · duration workflow-dependent · 720p open-weights TI2V · audio none in base text/image-to-video pipeline · verified 2026-07-11 - click any value to edit
{
  "action": "approaching strangers, quick interviews, reactions and walk-aways",
  "scene": {
    "location": "busy downtown street with heavy foot traffic",
    "mood": "chaotic street energy with unexpected encounters",
    "environment": "bustling city street with ambient urban sounds"
  },
  "camera": {
    "angle": "handheld street-level shots",
    "distance": "medium shots capturing both interviewer and subject",
    "movement": "mobile, following the action"
  },
  "lighting": {
    "type": "natural daylight with urban environment"
  },
  "style": {
    "tone": "energetic and unpredictable",
    "color_palette": "vibrant urban colors",
    "location": "{{location}}",
    "interview_question": "{{interview_question}}",
    "sample_answers": "{{sample_answers}}",
    "tone_style": "{{tone_style}}"
  }
}
  • Wan 2.2 is the latest version with genuinely open, downloadable weights.
  • Open-weights Wan 2.2 has no native audio track in the base text/image-to-video pipeline.
Runway · duration 2-10s · 720p; output dimensions vary by aspect ratio · audio none in video generation · verified 2026-07-11 - click any value to edit
{
  "action": "approaching strangers, quick interviews, reactions and walk-aways",
  "camera": {
    "angle": "handheld street-level shots",
    "distance": "medium shots capturing both interviewer and subject",
    "movement": "mobile, following the action"
  },
  "scene": {
    "location": "busy downtown street with heavy foot traffic",
    "mood": "chaotic street energy with unexpected encounters",
    "environment": "bustling city street with ambient urban sounds"
  },
  "lighting": {
    "type": "natural daylight with urban environment"
  },
  "style": {
    "tone": "energetic and unpredictable",
    "color_palette": "vibrant urban colors",
    "location": "{{location}}",
    "interview_question": "{{interview_question}}",
    "sample_answers": "{{sample_answers}}",
    "tone_style": "{{tone_style}}"
  }
}
  • Runway recommends clear, direct language; text-to-video should describe both visual elements and motion.
  • Image-to-video prompts should focus on describing the motion of the scene.

Fill in the variables

Type here (or click a suggestion) and every model tab above updates live.

{{location}} — Specific Location
Busy shopping district during weekendCollege campus between classesFood truck area during lunchOutside popular nightclub at closing timeSubway station during rush hourBeach boardwalk on summer day
{{interview_question}} — The Wild Question
What's the most illegal thing you've done that you'd do again?If you had to choose between your phone and your pet, what would you pick?What's a lie you tell yourself every day?If you could read minds for one day, whose mind would you avoid?What's something everyone does but nobody admits to?If dating apps showed people's biggest red flag, what would yours be?
{{sample_answers}} — Types of Responses Expected
Uncomfortable laughter followed by walking away quicklyOversharing personal details that get way too realAggressive defensiveness that reveals exactly what you suspectedCompletely unexpected wisdom that makes you question everythingHilarious misunderstanding of the questionPerfect one-liner that becomes the thumbnail quote
{{tone_style}} — Interview Vibe
provocativeplayfulconfrontationalcuriouschaotic

Run this prompt

Copy the JSON above, then paste it as your prompt in any of these tools. Veo 3.1 is free to try inside Gemini (daily limit applies).

More street interviews

all formats →