
An n8n AI video generator is a workflow that takes an idea in plain words and returns a finished video clip without anyone opening an editor. Ours starts from a chat message. It writes the scene, draws the opening frame with GPT Image, animates that frame into a 5-second vertical clip with Runway, records a one-line voiceover with ElevenLabs, and saves both files to Google Drive.
Below is the whole build, node by node, with the free JSON at the top. We also explain why the image is a JPEG and not a PNG, why the workflow asks Runway whether the clip is ready instead of waiting a fixed time, and what we changed from the version we started with.
What does the n8n AI video generator do?
It turns one sentence into one clip and one narration file. The canvas has six groups.
The first group is a chat trigger and one language model call. You type an idea, something like "a lighthouse keeper lighting the lamp as a storm rolls in", and the model returns three things as JSON: a prompt for the opening image, a prompt for the motion, and a line of narration.
The second group sends the image prompt to OpenAI's image API and gets back a portrait picture.
The third group is the Runway loop. It submits the picture and the motion prompt, then checks the job every 15 seconds until Runway reports that the clip has succeeded or failed.
The fourth group downloads the finished clip and uploads it to a Google Drive folder. The fifth sends the narration line to ElevenLabs and saves the MP3 to the same folder. The last group adds a row to a Google Sheet with the idea, the narration and both Drive links, then replies in the chat.
One thing it does not do: it does not lay the voiceover onto the video. You get two files with matching names, and you combine them in whatever editor you already use.
Download the n8n workflow JSON
The download is the full workflow, ready to import. It has every node, the notes that label each group, and the scene prompt in full.
Download the n8n AI video generator workflow (JSON)
To set it up:
- In n8n, go to Workflows, choose Import from File and pick the JSON.
- Add an OpenAI credential. Both the Write Scene node and the Generate Image node use it.
- Create two Header Auth credentials for the HTTP Request nodes. Runway uses the header
Authorizationwith the valueBearerfollowed by your API key, and it goes on Start Video and Check Video. ElevenLabs uses the headerxi-api-key, and it goes on Generate Narration. - Connect Google Drive and Google Sheets through their usual OAuth credentials.
- Replace the three placeholders:
YOUR_FOLDER_IDin Save Video and Save Narration,YOUR_SHEET_IDin Log to Sheet, andYOUR_VOICE_IDin the Generate Narration URL. - Give the sheet five column headers: created, idea, narration, videoLink and narrationLink.
- Open the chat, send one idea, and watch the first run through before you rely on it.
There are no keys in the file. We mention it because the export we started from had a live OpenAI key typed straight into an HTTP header field. n8n saves whatever you type into a node, so open any workflow JSON in a text editor before you send it to anyone.
What did we change from the original workflow?
The overall shape comes from an n8n course we worked through: a chat message, an image, a Runway clip, an ElevenLabs voice and a sheet. We rebuilt it as our own. The prompt is new, and the plumbing is different in eight places.
- It checks the video job instead of sleeping once. The original waited a fixed 90 seconds and then read the result a single time. If Runway was still working at second 91 there was no video to read, and the run failed at the very last step, after the voiceover had already been generated and paid for.
- It uses a Runway model that is on the current price list. The original called
gen3a_turbo, which no longer appears on Runway's API pricing page. Ours callsgen4_turbo. - The picture and the clip are the same shape. The original drew a portrait image and then asked Runway for a landscape video. Runway crops from the centre when the two do not match, so the top and bottom of every picture were thrown away.
- The image is labelled as what it is. OpenAI's image models return PNG unless you ask for something else, and the original handed that PNG to Runway marked as a JPEG. Ours asks OpenAI for a JPEG and labels it as one.
- One field name, spelled one way. The original read the prompt as
promptin one node andPromptin another. The model returns only one of them, so one of the two steps was working from a prompt that was not there. - Request bodies are built with
JSON.stringify. The original pasted the prompt into a hand-written JSON string, so a quotation mark or a line break in the model's answer broke the request. - The clip is copied to Google Drive. The original saved Runway's own link in the sheet, and those links expire.
- The voice reads a line written to be heard. The original sent the image prompt to the voice model, so the narration was a description of lighting and camera angles.
Why does one model call write three prompts?
Because the three tools need different things, and they have to agree with each other.
An image model wants a still: who is in the frame, where, in what light, and from which angle. A video model that starts from that picture can already see all of it. What it needs is the change, in one sentence: the keeper strikes a match, the camera pushes in. A voice model needs words that sound right when spoken, which rules out everything a visual prompt is made of.
So the Write Scene node asks for a JSON object with three keys, imagePrompt, motionPrompt and narration, and tells the model that all three must describe the same moment. The node has JSON output switched on, so each later step reads its own key and nothing has to be cut out of a block of text.
The prompt also sets limits that come from the tools. The image prompt stays under 900 characters. The motion prompt is one sentence with one subject movement and one camera move, because one clear action is easier for a video model to get right than three at once. The narration is capped at 30 words, which is about as much as fits in five seconds.
Why is the image a JPEG and not a PNG?
Size. The workflow passes the picture to Runway inline, as a base64 data URI, and Runway limits those to 5 MB after encoding. Base64 adds about a third to a file, so the picture itself has to stay under roughly 3.3 MB.
A detailed 1024 by 1536 PNG can go past that. The same picture as a JPEG at 90% quality is a fraction of the size, and the difference is invisible once it becomes the first frame of a moving clip. The Generate Image node asks for output_format: jpeg and output_compression: 90, and Start Video labels the data URI as image/jpeg to match.
The alternative is to upload the image somewhere public and pass Runway a URL. That works, but it adds a storage step and a file to clean up afterwards. For one frame, inline is simpler.
How does the workflow know when the Runway video is ready?
It asks. Runway's API is asynchronous: you submit a job, you get an id back straight away, and the clip turns up later. How much later depends on the model, the length and the queue.
Start Video sends the request and gets the task id. Wait 15 Seconds pauses, Check Video asks Runway for that task, and two If nodes read the status field. If it says SUCCEEDED, the run moves on to the download. If it says FAILED or CANCELED, the run stops with Runway's own failure message, so the execution log tells you why. Anything else means the job is still queued or running, and the loop goes back to the wait.
Both Runway nodes send the header X-Runway-Version: 2024-11-06, which the API requires on every call.
A job that never finishes and never fails would loop forever, so the workflow settings include a 30-minute execution timeout.
Why copy the video to Google Drive?
Because Runway's link is temporary. Runway's documentation on outputs says the URLs expire within 24 to 48 hours and tells you to download the file and store it yourself.
A sheet full of Runway links looks fine on the day and is full of dead links by the end of the week. So Download Video fetches the file as soon as the job succeeds, and Save Video uploads it to your Drive folder as clip- followed by the Runway task id. The narration is saved as narration- with the same id, which is how you tell which voice belongs to which clip.
What does each video cost to make?
The clip is most of it. On Runway's API price list, gen4_turbo costs 5 credits per second and a credit costs $0.01, so a 5-second clip is 25 credits, or $0.25. Changing duration to 10 in Start Video doubles that.
Each run makes four paid calls:
- Write Scene: OpenAI
gpt-4o-mini, a few hundred tokens. - Generate Image: OpenAI
gpt-image-1, one portrait image at medium quality. - Start Video: Runway
gen4_turbo, 25 credits ($0.25) for 5 seconds. - Generate Narration: ElevenLabs, about 30 words of speech.
The image and the narration are priced by OpenAI and ElevenLabs and depend on your plan and the quality you choose, so check their current rates. Prices change often. Work the total out against the number of clips you plan to make before you connect this to anything that runs on a schedule.
Can you use the videos as they are?
Sometimes. A five-second clip from a single image is good for a background loop, a social post or a storyboard frame. It is not a finished advert.
Expect to throw some away. Image-to-video models still get hands, text and fast motion wrong, and a prompt that reads well can produce a clip that does not. That is why the workflow ends in a sheet and a Drive folder instead of posting anywhere: a person looks at each clip before it is used. Our n8n YouTube Shorts generator works the same way, uploading every video as private, and our n8n LinkedIn post generator emails a draft instead of publishing.
If you'd rather have a pipeline like this built around your own product or content process, that's the work our AI automation service covers.
Common questions
Is the n8n workflow template free to download?
Yes. The JSON is free, there's no form in front of it, and there are no keys or credentials inside. Import it, connect your own accounts and change whatever you like.
Which AI models does the workflow use?
gpt-4o-mini writes the three prompts. OpenAI's gpt-image-1 draws the opening frame at 1024 by 1536. Runway's gen4_turbo turns that frame into a 5-second clip, and ElevenLabs' eleven_multilingual_v2 reads the narration. You can swap any of them in the node settings.
Do I need a paid Runway account?
You need Runway API credits, which are bought through Runway's developer portal and are separate from a subscription to the Runway app. Each 5-second clip uses 25 credits at the current rate.
Can n8n generate a video from text only, without an image?
This workflow always makes an image first, because gen4_turbo takes an image as its input. That is also a useful control: the picture fixes the look, and the motion prompt only has to describe what moves. Runway has other models that accept text alone, and you could call one from the same Start Video node.
How long can the video be?
The workflow asks for 5 seconds. You can raise duration in the Start Video node, and the cost rises with it. For anything longer than one shot, generate several clips and join them in an editor.
Why is the narration a separate file?
n8n has no built-in node that mixes audio into video. Keeping the MP3 next to the clip, with a matching name, lets you add it in your own editor or a rendering service. Our YouTube Shorts workflow shows one way to do that step with a template renderer.
Want something like this built and run in your own accounts? That's what our AI Automations service does.

