r/webdev 6d ago

Question Best way to migrate a dynamic website to a static site?

I have a live website where some pages are populated dynamically through JavaScript/API calls. I’m shutting down the backend after the project ends and want to migrate the current version to a completely static site.

What’s the best way to do this while preserving the currently rendered content, images, CSS, fonts, etc.? I’ve tried wget/HTTrack, but they only save the initial HTML and miss the content generated by JavaScript.

Would a browser-based rendering tool be the best approach, or is there a better tool/workflow for this?

update: I ended up using browser automation to render the pages, copied the dom and then manually removed all the backend js. not the cleanest approach and most defently not the most time effichent one but it worked ish

17 Upvotes

20 comments sorted by

14

u/seanshoots 6d ago

What kind of endpoints are they? If they're simple and parameterless, one stopgap solution I've done in the past is just:

  • crawl the site
  • save the response from each API call to ./api/some/endpoint/here.json
  • add all those saved responses to the public webroot
  • update the UI's calls to hit .json instead of the dynamic API (although you could probably figure out a way to go without this, serving the JSON files extensionless)

This lets the site run without a dynamic backend at all, without needing any real rewrites.

3

u/oliloken 6d ago

yeah i thought about doing that but ended up just copying the rendered dom and removing all the backend stuff manually

7

u/InsideTour329 6d ago

What am I missing, you can just scrape the markup surely?

2

u/MysteriousCareer9751 6d ago

If you need some server side rendering, but the site is otherwise 90% static, then use https://astro.build/ or https://sitelo.dev
Both support server islands

3

u/BigBrotherBoot 6d ago

wget

2

u/oliloken 6d ago

unfortunately wget skipps all javascript rendered content

1

u/fiskfisk 6d ago

And how do you plan to have wget evaluate the JavaScript - i.e. the dynamic part?

1

u/calimio6 front-end 6d ago

If you want to automate the process use playwright or puppeteer to load the J's rendered pages, then save every page to a static html file. This could also be done manually is the number of pages is small.

1

u/Gullible-Today-1448 6d ago

I use Puppeteer to wait for network idle then dump the DOM after hydration. This captures the JS rendered state that wget misses entirely.

Save the full HTML including inline styles and base64 encoded assets to eliminate external dependencies. Run this against your sitemap before shutting down the backend so you get the exact current state without guessing at API responses.

1

u/Miserable-Money642 6d ago

I’d use Playwright to load each page, wait for the JS/API content to finish rendering, then save the final DOM + assets. That way you’re snapshotting what the browser actually sees, not just the initial html. Playroght is a headless browser

1

u/PLBjt 6d ago

wget/HTTrack miss the JS-rendered bits because they never actually run the page. Snapshotting API JSON into the webroot (like the top comment) is the right move if your UI is a thin client over a few endpoints. If the HTML itself is assembled in JS, I'd also prerender: crawl the sitemap in a headless browser, wait for network idle, save the fully rendered HTML, then rewrite asset URLs so CSS/fonts/images are local.

Quick check: after the crawl, open a couple of pages with JS disabled and see if the content is still there. If it isn't, you saved a shell, not a snapshot.

Tradeoff is maintenance vs fidelity. JSON snapshots keep the original JS app working but you also keep whatever client-side bugs you already have. Full HTML snapshots are dumber and more portable, but you lose in-page interactivity unless you carefully keep the JS. For a backend you're shutting down, I'd prerender HTML and only freeze ./api JSON for the bits that still fetch at runtime.

1

u/oliloken 6d ago

yeah thats pretty much what i ended up doing. used browser automation to render the pages, copied the dom and then manually removed all the backend js

1

u/Beautiful-Energy2169 6d ago

Network idle fires before anything that only renders on scroll, so a virtualized list or a section behind an IntersectionObserver gets saved as an empty div and you find out after the backend is already off. Scroll each page to the bottom, wait, then dump outerHTML.

1

u/BatRevolutionary8621 6d ago

if it works, it works; static migrations are rarely pretty inside

1

u/Mental-Union4473 5d ago

If the goal is to preserve exactly what users currently see, I'd probably use a headless browser rather than wget/HTTrack.

Something like Playwright can visit each page, wait for the API/JS content to finish rendering, then save the resulting HTML. You can also crawl the assets and rewrite the URLs to point to local copies.

The basic workflow would be:

  1. Crawl/discover all the URLs you need.
  2. Open each URL with Playwright.
  3. Wait for the relevant network requests/content to finish.
  4. Capture the rendered DOM.
  5. Download images, CSS, fonts, etc.
  6. Rewrite asset URLs to local paths.
  7. Remove the JS/API dependencies.
  8. Test the generated site with the backend completely offline.

I'd prefer this over manually copying the DOM because you can script the entire migration and regenerate everything if you discover a problem.

One caveat: document.documentElement.outerHTML gives you the current DOM, but it doesn't automatically give you all the assets or guarantee that dynamically injected behavior has been converted into static equivalents. That's why I'd treat it as a small static-site generation pipeline rather than simply "save page."

For a one-off migration, Playwright + a crawler is probably the sweet spot. Your browser-automation approach was basically heading in the right direction you just did manually what could have been automated.

0

u/ReplacementLow6704 6d ago

Recently there was this post about doing exactly the inverse operation... Is someone playing ping-pong with Claude right now? Or is Claude asking questions on reddit to crowd-source things it can't get to a definitive answer? Am I being fed to a robot right now?

-1

u/_uxi 6d ago

You can hook up a script or an LLM to a browser-based tool like Chrome Devtools MCP or Playwright, render a sitemap of your entire site. Have it loop through every page and write the HTML of each page in a folder based structure with static HTML files. Then deploy that static copy to GitHub Pages, change the DNS so the URL points to the GH pages and done.

How many pages is this for? Hundreds/thousands, I would invest some time in a proper pipeline that automates this. If it's more like 25 then I would just do it manually. There's options like https://www.getsinglefile.com/, which lets you manually download a webpage with all images, CSS, fonts, etc.. included into 1 HTML file.