Skip to content

Website migration

Crawl a public website, see every image and file it references, import the ones you want into a bucket, and get a map from old URLs to stable SteadyLink links.

On this page

Website migration is for moving an existing site's media into SteadyLink without downloading and re-uploading it by hand. You give SteadyLink a starting URL; it crawls up to 200 pages on that site, lists every image, video, document, and download those pages reference, and checks whether each one still loads. You pick the files to keep, and an import job copies them into a bucket folder. The result is a CSV or JSON map from each original URL to its new stable link, which you use to update your site's source.

Before you start#

  • The site must be reachable on the public internet over HTTP or HTTPS on the standard ports (80 or 443). Addresses that resolve to private, loopback, link-local, multicast, or reserved ranges, localhost, and cloud metadata hosts are refused, including through redirects. A URL with a username or password in it is refused too.
  • The crawler identifies itself as SteadyLink-Migration-Scanner/1.0 and follows the site's robots.txt. Pages disallowed for that user agent are skipped. If your staging site blocks crawlers, allow this user agent for the duration of the scan.
  • Decide which bucket and folder the files should land in. You can type a new folder path; it is created during the import.

Run a scan#

  1. Start the scan

    Open Migrate a site, enter a domain or starting URL such as northwind.example or https://northwind.example/products, and start the scan. A bare domain is treated as https://. The dashboard scans up to 50 pages; use the API to scan up to 200.

  2. Wait for the report

    The scan runs in the background. Progress, pages scanned, and media found update as it goes. You can leave the page and come back through History.

  3. Review the report

    When the status is complete, the report lists every media URL found, with its type, size, status, and how many pages reference it.

Start a scan with the API
curl -X POST https://api.steadylink.io/api/migration-scans \
  -H "X-API-Key: $STEADYLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://northwind.example", "maxPages": 200 }'
202 AcceptedResponse
{
  "id": "a3c5e7f9-1b2d-4e6f-8a0c-2e4f6a8b0c1d",
  "url": "https://northwind.example",
  "status": "queued",
  "maxPages": 200,
  "pagesScanned": 0,
  "mediaFound": 0,
  "progress": 0,
  "error": null,
  "createdAt": "2026-10-08T15:20:44.019311",
  "updatedAt": "2026-10-08T15:20:44.019311"
}

Poll GET /api/migration-scans/{scan_id} until status is complete or failed. maxPages accepts 1 to 200 and defaults to 50.

How the crawler explores#

The crawl starts at your URL and follows links breadth-first, staying on the exact host you started on (www.northwind.example and northwind.example are different hosts). It stops when it has visited maxPages pages or runs out of links. Each page can be up to 3 MB; larger pages and pages that fail to load are skipped and do not stop the scan. Only HTML responses are read for references; other content types are skipped.

Media can live on any public host, such as a CDN or an old asset domain. The crawler records up to 5,000 distinct media URLs per scan.

What counts as media#

ReferenceRecorded as
<img>, <source>, <video>, <audio> src, poster, and every srcset candidateimg, source, video, audio
<link rel="icon"> and <link rel="preload">link
og:image, og:video, og:audio, and twitter:image meta tagsopen_graph
url(...) in an element's inline style attributecss_background
Links to files ending in a media or document extension, or links with a download attributedownload

Download extensions are .avif, .bmp, .csv, .doc, .docx, .gif, .ico, .jpeg, .jpg, .mp3, .mp4, .pdf, .png, .ppt, .pptx, .svg, .tif, .tiff, .webm, .webp, .xls, .xlsx, and .zip. URLs starting with data:, javascript:, mailto:, or tel: are ignored. Background images declared in external stylesheets are not followed.

Read the report#

Each media entry in the report has:

FieldMeaning
originalUrlThe absolute URL as referenced, without its fragment.
mediaTypeWhere it was found, from the table above.
contentType, byteSizeWhat the server reported when probed. The scan does not download whole files, so this comes from response headers.
statusCodeThe HTTP status of the probe.
brokentrue when the URL failed to load or returned 400 or higher.
pageUrls, referencesThe pages that reference it, and how many. The report lists the most referenced files first.
importedAssetId, newUrlFilled in once the file is imported.

A broken entry is worth fixing on your current site whether or not you migrate: it is a reference that already fails for your visitors.

Import the files you want#

Select files in the report, choose a Destination bucket and Folder (default website-import), and choose Import n selected. Through the API:

Terminal
curl -X POST https://api.steadylink.io/api/migration-scans/a3c5e7f9-1b2d-4e6f-8a0c-2e4f6a8b0c1d/imports \
  -H "X-API-Key: $STEADYLINK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "mediaIds": ["4d6f8a0c-2e4f-4a8b-9c1d-3e5f7a9b1c2d", "6f8a0c2e-4f6a-4b0c-8d2e-5f7a9b1c3d4e"],
    "bucketId": "0c9e4b1a-6d2f-4a7e-b3c5-8f1d2e3a4b5c",
    "path": "website-import/images"
  }'
202 AcceptedResponse
{ "queued": true }

One import can include up to 200 files. Every selected file must belong to this scan, must not be marked broken, and must not already be imported; otherwise the request fails with 422 "Scan, bucket, or selected media is invalid". The scan itself must be complete. The folder path can be nested; . and .. segments and control characters are rejected.

While the import runs, the scan's status is importing. For each file, SteadyLink:

  1. Downloads it from the original URL, up to your plan's upload size limit.
  2. Detects the real type from the bytes. A file whose bytes do not match the type its server declared (for example an HTML error page served as image/png) is skipped and marked broken instead of being imported.
  3. Scans it for malware and saves it into the folder. File names come from the URL path; if a name is already taken in that folder, a suffix is added (hero.jpg, hero-2.jpg) so nothing is overwritten.
  4. Records the new asset and its stable link, https://cdn.steadylink.io/a/{asset_id}, on the report entry.

Files that fail any step are marked broken and the import continues with the rest. When it finishes, the scan returns to complete. If a later scan finds a URL you already imported into the same bucket folder, importing it again reuses that file instead of creating a duplicate.

Export the URL map#

Once files are imported, download the map and use it to update references in your site.

  • In the dashboard, choose Export map on the report to download a CSV.
  • Through the API, call GET /api/migration-scans/{scan_id}/export/csv or /export/json. Only imported files are included.
steadylink-mapping-a3c5e7f9-1b2d-4e6f-8a0c-2e4f6a8b0c1d.csv
originalUrl,steadyLinkUrl
https://northwind.example/wp-content/uploads/2024/03/hero.jpg,https://cdn.steadylink.io/a/5b8e1f2a-3c4d-4e6f-9a0b-1c2d3e4f5a6b
https://cdn.oldhost.example/brochure.pdf,https://cdn.steadylink.io/a/9c0d1e2f-3a4b-4c5d-8e6f-7a8b9c0d1e2f

The JSON export contains the same mapping array plus instructions and short examples for HTML, Next.js, and CSS. Search for each original URL in your source, replace it, review the diff, and deploy as you normally would. Remember that srcset attributes and CSS files may reference the same image several times.

Once your site points at SteadyLink links, you can update any of those files later by replacing it; the links in your site never need to change again.

Cancel or retry#

  • Cancel with POST /api/migration-scans/{scan_id}/cancel while a scan is queued, running, or importing. Work stops at the next page or file. Files already imported stay in your bucket.
  • Retry a scan whose crawl failed with POST /api/migration-scans/{scan_id}/retry. It starts the crawl again from the beginning.
  • Resume a failed import by sending the import request again for the files that are not yet imported. If an import fails partway, the scan's error starts with "Import failed:" and the import route accepts the scan even though it is not complete.

Limits#

LimitValue
Pages per scan1 to 200
Files per import request200
Media URLs recorded per scan5,000
Scans started per workspace20 in any 24 hours, then 429
Scanned pages per billing periodFree 100, Personal 1,000, Pro 10,000, Business 100,000, Enterprise 1,000,000

Pages count against the monthly allowance when the scan starts: a queued or running scan reserves its full maxPages, and a finished scan counts the pages it actually visited. Asking for more pages than remain returns 409 with the code feature_limit_reached, so lower maxPages or wait for the next billing period.

Starting scans and imports requires a workspace role that can edit files (member, admin, or owner) or an API key with assets:write. Reading reports and exports needs assets:read.

Next steps#