- Published on
The Server Never Touches the Video: IKEA Mexico's First AI-Powered Campaign
Overview
The stakes make it worth telling. IKEA is still relatively new to the Mexican market, and Ofrenda was its first campaign that both ran on its own website and handed each visitor a generated, personalized asset: a family photo turned into an animated Día de Muertos keepsake, seconds after upload. There is no template for that at IKEA Mexico yet. We were writing one.
The whole edit lives inside a URL: a Cloudinary transformation string that describes the composite, and nothing else. The CDN does all the video processing, so none of it ever touches our servers, which is the whole point, because this was a campaign built for massive traffic. This is the story of how that string got written, and why I now believe the hard part of AI video isn't generating pixels. It's everything around them.


And the deliverable itself, playing on loop: one photo, animated, composited into the ofrenda frame.
Why It Matters
Calling a video model is no longer a skill. Anyone with an API key can generate five seconds of moving pixels in an afternoon. What separates a demo from a product is the last 10%: the brand compliance, the aspect ratios, the cost caps, the iteration speed. All the unglamorous constraints that a campaign lives or dies by.
The URL-as-pipeline idea solved the biggest one. In a traditional motion-graphics workflow, every change means a re-export: open the timeline, nudge the overlay, render, upload, store the new artifact. Multiply that by every client revision and every A/B variation. It's v7_FINAL_final.mp4 all the way down.
With transformation URLs, the edit is a pure function. The generated video is stored exactly once. Want the overlay 8% smaller? Change overlay_width from 235 to 218 and you have a new URL. New frame, new font, new text position — different arguments, same asset, zero re-renders. A parameter change is a deploy, and every composite is processed at the CDN's edge, on their compute, not ours. For a massive campaign, that difference is the architecture: no transcoding fleet to size, scale, or pay for, no matter how many visitors show up.
Tech Stack
- FastAPI on Vercel serverless Python: the conductor
- Gemini 2.0 Flash via a pydantic-ai agent for validation
- nano-banana (Replicate) to retouch the portraits
- wan-2.2-i2v-fast (Replicate) to generate the motion
- Cloudinary transformations as the compositor
- Supabase Postgres + Storage for the memory
What It Does
The pipeline is a relay race, and each model runs one leg.
Gemini is the bouncer
Every upload hits a Gemini agent with a strict-JSON contract: is_valid_public, subject_type, is_portrait. Humans and pets only, no nulls, no prose. I wrote the prompt the way everyone writes those prompts: "Respond with ONLY a single JSON object and NOTHING else."
The model did not always comply.
So underneath the prompt sits _safe_json_parse, a function whose entire job is to scrape the first {...} block out of whatever the model actually said. Prompting is negotiation; parsing is law. I learned this building a chatbot that could never break character, and the lesson transfers perfectly: the prompt asks nicely, the code enforces.
nano-banana is the retoucher
Valid portraits get conditioned by nano-banana before animation. The interesting part is one line in the prompt: "crop/zoom only; do NOT alter identity or features."
When the input is a photo of someone's late grandmother, that stops being a prompt-engineering flourish and becomes the ethics policy of the entire product. The model is allowed to reframe the face. It is not allowed to improve it. Identity preservation as an engineering constraint, not a disclaimer buried in terms of service.
wan-2.2 is the actor
The retouched portrait goes to wan-2.2-i2v-fast: 81 frames, 16 fps, 480p, with the prompt "he dances in peace, in slow motion." Every generated video in this campaign has that same gentle motion, because the shelf it lives in is an ofrenda, and the movement has to feel like a memory, not a deepfake.
Cloudinary is the editor
The generated clip uploads to Cloudinary, and then the edit happens, as text. One transformation chain says: overlay this video onto the animated Kallax frame, force 2:3 aspect ratio, fill from center, round the corners 13 pixels, stamp the family member's name in Montserrat at 30pt in #0058A3, pick a safe codec.
That aspect ratio was a war. The overlay kept coming back distorted: faces stretched, shoulders wrong. Every prompt tweak on the generation side did nothing, because the distortion was happening downstream, in the composition. c_scale resized without caring about shape. The fix was two parameters: ar_2:3 plus c_fill with center gravity. Crop the overflow, keep the proportions. There's a commit in the history that is literally just fix(cloudinary): enforce 2:3 AR using ar_2:3 + c_fill, and it's one of my favorites, because it fixed a customer's grandmother.
The exact transformation string
The full compose_video_with_center_overlay chain we ship, parameter by parameter, plus the debug recipe I use for aspect-ratio drift (which layer to suspect first, and why).
Supabase is the memory
Users, consent (accept_mkt), and every generated video land in Supabase, with row-level security so nobody reads somebody else's grandmother. And when storage fails (serverless storage occasionally does), the endpoint streams the raw MP4 straight back to the customer anyway. Never send someone away empty-handed from a memorial.
Impact / Lessons
What stuck with me:
- The edit lives in the URL, not the file. I stopped thinking of compositions as artifacts. They're expressions now: arguments against stored assets, evaluated lazily by a CDN.
- Rate limits are product design. A low requests-per-minute limit is as much a cost cap as it is abuse prevention. The invoice math for an AI activation gets decided in the router decorator, not in the cloud bill review.
The rate-limit math
How to cost-cap an AI activation before the invoice caps it for you: the per-request cost breakdown across the three model calls, and how I chose the limiter values.
- Brand guidelines will out-vote your taste. I built a whole endpoint for uploading custom fonts, tested a romantic script face (Parisienne) for the name overlay… and one commit later,
fix(overlay): swap Parisienne with Montserratreplaced it. IKEA blue is#0058A3or it isn't IKEA. I stopped grieving quickly. - The moat is the last 10%. Generating the video took an afternoon. Making it correct took everything else: validated, identity-safe, aspect-true, brand-exact, capped, and resilient to Supabase having a bad day.
The most useful shift in my own head: I used to ask which model should generate this? Now I ask where exactly does the model's job end and the code's begin? Draw that line early, enforce it server-side, and the pipeline survives contact with real customers.
Built with FastAPI, Replicate (nano-banana, wan-2.2-i2v-fast), Gemini 2.0 Flash, Cloudinary, and Supabase.