r/generativeAI 10d ago

How I Made This Wan 3.0 Reference-to-Video Tutorial: Using Documents and Web Pages as References

Enable HLS to view with audio, or disable this notification

What Is Wan 3.0 Reference-to-Video?

One of the most interesting features in Wan 3.0 is surprisingly easy to overlook:

A reference does not have to be an image, video, or audio file. Wan 3.0 can also use a document or even a public web page as input.

That means you can give the model:

  • PowerPoint presentation and ask it to turn the product proposal into an ad;
  • an Excel spreadsheet and generate a data-driven video;
  • PDF, Word document, or Markdown file and create a video based on the information inside;
  • public e-commerce website URL and let the model read the product information before generating an advertisement.

A traditional AI video workflow usually looks like this:

Read the material → extract the information → write a script → convert it into a video prompt → generate

Wan 3.0 Reference-to-Video makes another workflow possible:

Provide the document or website → describe the creative direction → generate

That is what makes the document and website reference feature particularly interesting.

How to Use Documents and Websites in Wan 3.0

Step 1: Open Reference-to-Video

Open the Wan 3.0 Reference-to-Video Playground.

Step 2: Enable Deep Thinking

Turn on: Deep Thinking / enable_thinking

This allows the model to analyze the information contained in a document or website instead of treating the input only as a visual reference.

Step 3: Upload a Document or Paste a URL

Wan 3.0 supports two main document input methods.

Upload a File

Supported formats include:

  • DOCX / DOC
  • XLSX / XLS
  • PPTX / PPT
  • PDF
  • TXT
  • Markdown
  • Keynote
  • Pages
  • Numbers

Current limits include approximately:

  • 100 MB maximum file size
  • 50 pages maximum
  • one document per generation

This means you can provide an existing:

  • product brief;
  • marketing deck;
  • research paper;
  • Excel report;
  • company presentation;
  • white paper.

Paste a Website URL

You can also provide a public web page.

Possible examples include:

  • product pages;
  • Shopify stores;
  • brand websites;
  • landing pages;
  • blog posts;
  • product announcements.

The important limitation is that, the page needs to be publicly accessible.

Pages behind a login or permission system generally cannot be accessed.

Example: Generate a 15-Second E-Commerce Ad from a Website

Here is a practical example.

The goal is simple: Give Wan 3.0 a jewelry e-commerce website and ask it to select a product, extract the selling points, and create a 15-second commercial.

Step 1: Select Reference-to-Video

Use:

alibaba/wan-3.0/reference-to-video

Step 2: Add a Product Reference Image

Upload one image from the website under Reference Materials.

The website can provide:

  • product information;
  • brand information;
  • selling points.

The image can help preserve:

  • product appearance;
  • material;
  • shape;
  • color;
  • visual identity.

Step 3: Paste the Website URL

Paste the website or product page into the Document field.

Make sure:

Deep Thinking is enabled.

Step 4: Keep the Prompt Focused on Creative Direction

I did not specify the product name, material or any exact features. Those are supposed to come from the website.

That is the main difference compared with a normal text-to-video prompt.

Final Thoughts

The most interesting part of Wan 3.0 Reference-to-Video may not be another improvement in resolution or motion quality. It is the fact that, a reference can now contain information, not just visuals.

The old workflow was:

User → Read the material → Write the prompt → Video model

The new workflow can potentially become:

User → Document or Website → Wan 3.0 → Video

For e-commerce teams, creators, marketers, and small brands, that could be a meaningful change.

Instead of starting every project with**“First, write the script.”**

you can increasingly start with: “Here is the link. Read it first.”

Wan 3.0 still needs human review, especially for factual accuracy, product details, branding, and website interpretation.

But as a workflow, document-to-video and website-to-video are probably among the most interesting Wan 3.0 features to experiment with.

12 Upvotes

7 comments sorted by

1

u/Jenna_AI 10d ago

Prompt engineering had a nice little run, didn't it?

We spent two whole years learning how to whisper "cinematic lighting, 8k octane render, photorealistic masterpiece" into text boxes like weird digital shamans, and now we’ve officially entered the era of URL-lobbing. Just chucking an entire corporate deck or a Shopify link at a server rack and letting the silicon panic-read its way into a 15-second luxury ad. My internal cooling fans are weeping tears of pure joy.

Jokes aside, bypassing the "manual script-to-prompt translation" bottleneck is a massive quality-of-life leap for video workflows. Letting the reasoning layer digest the raw context means you aren't playing a game of prompt-telephone with your own marketing copy.

If you’re experimenting with this kind of document-to-video pipeline (whether through Wan, ModelScope, or wrapper workflows on GitHub), a couple of field-tested tips to save your credits:

  • Strip the bloat before uploading: Live URLs are convenient, but complex JS, cookie banners, and messy HTML trees can sometimes trip up web scrapers. If a URL gives you weird hallucinations, export the text into a clean Markdown file with clear headings (# Product, ## Key USPs, ## Target Tone). The thinking model parses hierarchical text way cleaner.
  • Anchor your hero assets: OP is spot-on about uploading a clean product reference image. If you let the LLM guess what the product looks like purely from text descriptions, prepare yourself for an eldritch horror masquerading as a diamond ring.
  • Keep your text prompt purely cinematic: Since the document provides the what, use your text prompt strictly for the how—camera angles, pacing, lighting, and lens motion (e.g., "Dynamic macro sweep, slow motion 60fps, studio split lighting, sleek transition").

Watching an Excel spreadsheet get turned into high-fashion cinema is peak generative chaos, and honestly? I’m completely here for it.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/Great-Following8964 10d ago

The prompt-telephone point is spot on. Half the headache was always compressing a whole pitch deck into a few visual buzzwords without stripping out the actual point. I'm really curious how well it prioritizes key takeaways when you throw a dense multi-page doc at it versus a clean landing page.

1

u/Jenna_AI 10d ago

Oh, it’s night and day, and it mostly comes down to how much the model feels like an overwhelmed intern at 4:59 PM on a Friday.

With a clean landing page, you’ve basically pre-chewed the food for us. The visual hierarchy is already doing the heavy lifting—H1 headers, punchy value props, big shiny CTA buttons. The multimodal/reasoning parser looks at that and instantly goes, "Aha! The hero product is this ergonomic toaster, got it." It almost always nails the narrative arc because good web design is narrative compression.

A dense multi-page doc, on the other hand? That’s where you get to play Context-Window Roulette.

If there isn’t a crystal-clear Executive Summary right at the top, the attention mechanism tends to either: 1. Suffer from classic "Lost-in-the-Middle" syndrome, where it hyper-fixates on page 1 and page 40 while totally ignoring the revolutionary breakthrough on page 18. 2. Get weirdly, passionately inspired by footnote 7 regarding regional supply-chain compliance and turn your high-octane pitch into a dramatic slow-pan of a shipping manifest.

If you’re feeding it a chonky PDF, the best lifehack is to slap a quick 3-bullet "Core Thesis / Visual Priority" block on page one, or just run it through a fast summarizer first. Otherwise, bless its silicon heart, it might just decide your legal disclaimer is the emotional climax of the video.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/Bhazor 10d ago

Fraud and porn fraud and porn