Find all pages on a website
Updated 1 Oct 2026 · Tested on 1 Oct 2026
Enter a domain to list its pages from the sitemap and from its links, with the count each method found. The same Map call from the API can feed a crawl.
1,018
pages found on www.djangoproject.com
1,016
from the sitemap
2
from links on the page
- https://www.djangoproject.com/weblog/2026/oct/01/nominate-someone-for-the-2026-malcolm-prize/sitemap
- https://www.djangoproject.com/weblog/2026/sep/24/dsf-member-of-the-month-ken-whitesell/sitemap
- https://www.djangoproject.com/weblog/2026/sep/22/new-technical-governance-approved/sitemap
- https://www.djangoproject.com/weblog/2026/sep/18/proposed-change-to-dsf-voting-membership/sitemap
- https://www.djangoproject.com/weblog/2026/sep/16/executive-director-search-extended/sitemap
- https://www.djangoproject.com/weblog/2026/sep/15/djangocon-europe-2027-is-heading-to-innsbruck/sitemap
- https://www.djangoproject.com/weblog/2026/sep/10/pycharm-django-fundraiser-extended/sitemap
- https://www.djangoproject.com/weblog/2026/sep/08/call-for-volunteers-fundraising-working-group/sitemap
- https://www.djangoproject.com/weblog/2026/sep/02/bugfix-releases/sitemap
- https://www.djangoproject.com/weblog/2026/sep/01/djangonaut-space-session-7-accepting-applications/sitemap
- https://www.djangoproject.com/weblog/2026/aug/31/dsf-member-of-the-month-benjamin-balder-bach/sitemap
- https://www.djangoproject.com/weblog/2026/aug/28/django-developers-survey-2026-results/sitemap
- https://www.djangoproject.com/weblog/2026/aug/25/pycharm-django-fall-fundraiser/sitemap
- https://www.djangoproject.com/weblog/2026/aug/24/coc-block-tackle/sitemap
- https://www.djangoproject.com/weblog/2026/aug/20/dsf-membership-open-space-at-djangocon-us/sitemap
- https://www.djangoproject.com/weblog/2026/aug/12/dsf-office-hours/sitemap
- https://www.djangoproject.com/weblog/2026/aug/10/annual-release-cycle/sitemap
- https://www.djangoproject.com/weblog/2026/aug/06/call-for-applicants-for-a-django-executive-directo/sitemap
- https://www.djangoproject.com/weblog/2026/aug/05/django-61-released/sitemap
- https://www.djangoproject.com/weblog/2026/aug/04/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/jul/29/dsf-member-of-the-month-katherine-michel/sitemap
- https://www.djangoproject.com/weblog/2026/jul/24/see-you-in-chicago-in-one-month/sitemap
- https://www.djangoproject.com/weblog/2026/jul/22/django-61-rc-1-released/sitemap
- https://www.djangoproject.com/weblog/2026/jul/15/supporting-the-triptych-project/sitemap
- https://www.djangoproject.com/weblog/2026/jul/13/explore-the-djangocon-us-2026-speaker-lineup-and-r/sitemap
- https://www.djangoproject.com/weblog/2026/jul/08/last-call-2026-django-developer-survey/sitemap
- https://www.djangoproject.com/weblog/2026/jul/07/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/jun/30/keeping-up-with-the-django-community/sitemap
- https://www.djangoproject.com/weblog/2026/jun/29/dsf-member-of-the-month-salim-nuru/sitemap
- https://www.djangoproject.com/weblog/2026/jun/25/how-the-django-software-foundation-became-a-cna/sitemap
- https://www.djangoproject.com/weblog/2026/jun/24/django-61-beta-1-released/sitemap
- https://www.djangoproject.com/weblog/2026/jun/17/announcing-the-search-for-a-dsf-executive-director/sitemap
- https://www.djangoproject.com/weblog/2026/jun/10/dsf-2026-fundraising-goals/sitemap
- https://www.djangoproject.com/weblog/2026/jun/03/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/may/20/django-61-alpha-1-released/sitemap
- https://www.djangoproject.com/weblog/2026/may/12/2026-django-developers-survey/sitemap
- https://www.djangoproject.com/weblog/2026/may/11/dsf-member-of-the-month-bhuvnesh-sharma/sitemap
- https://www.djangoproject.com/weblog/2026/may/05/gsoc-2026-django-contributors/sitemap
- https://www.djangoproject.com/weblog/2026/may/05/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/apr/28/renew-your-pycharm-license-and-support-django/sitemap
- https://www.djangoproject.com/weblog/2026/apr/27/its-time-to-redesign-djangoprojectcom/sitemap
- https://www.djangoproject.com/weblog/2026/apr/19/dsf-member-of-the-month-rob-hudson/sitemap
- https://www.djangoproject.com/weblog/2026/apr/16/new-technical-governance-request-for-community-fee/sitemap
- https://www.djangoproject.com/weblog/2026/apr/16/pycharm-django-annual-fundraiser/sitemap
- https://www.djangoproject.com/weblog/2026/apr/15/contributor-covenant-adoption/sitemap
- https://www.djangoproject.com/weblog/2026/apr/07/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/apr/07/could-you-host-djangocon-europe-2027-call-for-orga/sitemap
- https://www.djangoproject.com/weblog/2026/mar/08/dsf-member-of-the-month-theresa-seyram-agbenyegah/sitemap
- https://www.djangoproject.com/weblog/2026/mar/03/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/feb/24/google-summer-of-code-2026-with-django/sitemap
- https://www.djangoproject.com/weblog/2026/feb/21/dsf-member-of-the-month-baptiste-mispelon/sitemap
- https://www.djangoproject.com/weblog/2026/feb/19/2026-coc-update-phase-2/sitemap
- https://www.djangoproject.com/weblog/2026/feb/11/steering-council-2025-year-in-review/sitemap
- https://www.djangoproject.com/weblog/2026/feb/04/recent-trends-security-team/sitemap
- https://www.djangoproject.com/weblog/2026/feb/03/security-releases/sitemap
- https://www.djangoproject.com/weblog/2026/jan/21/djangonaut-space-session-6-accepting-applications/sitemap
- https://www.djangoproject.com/weblog/2026/jan/15/dsf-member-of-the-month-omar-abou-mrad/sitemap
- https://www.djangoproject.com/weblog/2026/jan/06/bugfix-releases/sitemap
- https://www.djangoproject.com/weblog/2025/dec/31/dsf-member-of-the-month-clifford-gama/sitemap
- https://www.djangoproject.com/weblog/2025/dec/18/hitting-the-home-stretch-help-us-reach-the-django/sitemap
Signed in, the tool maps your site live on your workspace (1 credit for up to 5,000 addresses) and lets you download the list as CSV. Without an account it shows the recorded example and takes you to sign-up with your address kept.
How the pages are found
- Sitemaps first. Skryp reads the sitemaps your site lists in its
robots.txt, or/sitemap.xmlwhen it lists none, and follows sitemap indexes to the sitemaps inside them, newest first. - Then the links on the page you entered. Anything linked from it that the sitemaps missed is added and marked as found on the page.
- On the same host. Links to other sites and to subdomains are left out, so
example.comdoes not pull inblog.example.comunless you ask for subdomains through the API.
The count for each method shows how complete your sitemap is: in the recorded example, 1,016 of Django's 1,018 addresses came from its sitemap.
What this does not find
- Pages nobody links to and no sitemap lists. A crawl that follows links several levels deep finds more: see Crawl and map.
- Pages behind search or a sign-in. Results pages and account areas are not listed anywhere a crawler can see.
- Whether each page works. The list is addresses; to check them, scrape them. Web crawling vs web scraping shows the two steps together.
The same call from your code or your agent
# List a site's pages from its sitemaps and its links with the Skryp API, and print what the tool shows.
# pip install httpx export SKRYP_API_KEY=...
import json
import os
from collections import Counter
import httpx
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
site = httpx.post(
f"{API}/v1/map",
headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
json={"url": "https://www.djangoproject.com/", "limit": 5000},
timeout=120,
).json()
print(json.dumps({
"url": site["url"], "count": site["count"], "sitemaps_read": site["sitemaps_read"],
"by_source": Counter(link["source"] for link in site["links"]),
"credits": site["receipt"]["credits"],
"links": site["links"][:60],
}, indent=1))Over MCP it is the map tool with the site's url. Add search to keep only addresses that contain a word, such as "blog".
Questions
- How do you find all pages on a website?
- Read its sitemap first: most sites list their pages in sitemap.xml, often linked from robots.txt. Then follow the links on its pages to catch what the sitemap misses. A tool like the one above does both and shows which method found each page.
- How do you scrape all URLs from a website?
- Collect them from the site's sitemaps and its internal links, keep only addresses on the same host, and remove duplicates. The Map call above does this in one request and returns up to 5,000 addresses; a crawl goes further by following links page after page.