Sitemap

NAME

Sitemap - Sitemap generator for Raku

SYNOPSIS


use Sitemap;

# Generate XML from URL list
my $builder = Sitemap::Builder.new;
$builder.add-item('https://example.com/page1', :lastmod(DateTime.now), :priority(0.8));
$builder.add-item('https://example.com/page2');
$builder.write: 'sitemap.xml';

# Or crawl a website
my $crawler = Sitemap::Crawler.new(url => 'https://example.com');
$crawler.on-add: -> $url { say "Found: $url" };
$crawler.on-done: -> @urls { say "Total: @urls.elems()" };
$crawler.start;

# Fetch sitemap from server (with auto-discovery)
use Sitemap::Fetcher;
my $xml = fetch('https://example.com', :ssl-verify(False));
# Recursive fetch for sitemap indexes (downloads all child sitemaps)
my %result = fetch-recursive(
    'https://example.com/sitemap.xml',
    :format<xml>,
    :verbose
);

DESCRIPTION

Sitemap is a Raku module for generating XML sitemaps, crawling websites, and fetching sitemaps from servers. It supports:

  • Generate standard XML sitemaps (sitemap.org protocol)

  • Sitemap index files for large sites (>50k URLs)

  • Parse existing sitemaps (including recursive parsing of indexes)

  • Fetch sitemaps from remote servers (with recursive support)

  • Automatic gzip decompression of .gz sitemaps

  • Automatic sitemap discovery (robots.txt + /sitemap.xml fallback)

  • Recursive fetch creates descriptive directories (e.g., example-com-sitemap/)

  • Child sitemaps named with path context (e.g., sitemap-blog.xml)

  • Build hierarchical site trees with depth-inferred priorities

  • Crawl websites to discover URLs

  • Respect robots.txt rules

  • Output in multiple formats (XML, HTML, TXT, RSS, Atom, mRSS)

Sitemap::SiteTree

Define your site as a tree of URL stubs with automatic priority computation:


use Sitemap::SiteTree;

my $root = Sitemap::SiteTree.new;
my $blog = $root.add-child('blog');
$blog.add-child('first-post');
$blog.add-child('second-post');

my $builder = $root.to-builder('https://example.com');
$builder.write: 'sitemap.xml';  # priorities: 1.0, 0.8, 0.6, 0.6

Also supports flat-list wiring via wire-parents and JSON import via the sitemap tree CLI command.

Sitemap::Fetcher

The Sitemap::Fetcher module provides functions for fetching sitemaps:

  • fetch($url, :$ssl-verify, :$verbose) - Fetch a single sitemap with automatic gzip decompression

  • fetch-recursive($url, :$format, :$output, :$force, :$ssl-verify, :$verbose, :$concurrency, :$follow-foreign-children, :$max-children) - Fetch sitemap index and all child sitemaps

  • discover-sitemap($domain, :$ssl-verify, :$verbose) - Auto-discover sitemap (robots.txt + /sitemap.xml fallback)

  • output-filename($input, :$format, :$output-override) - Generate filename with path context

  • format-output(@items, $fmt, $title, $link, :$pretty) - Format items to various output formats

AUTHOR

Sasha Abbott

COPYRIGHT AND LICENSE

Copyright 2026 Sasha Abbott

This library is free software; you can redistribute it and/or modify it under the CC0 1.0 Universal license.

Sitemap v0.0.1

Sitemap generator for Raku

Authors

  • Sasha Abbott

License

CC0-1.0

Dependencies

Cro::HTTP:auth<zef:cro>:api<0>URI:auth<zef:raku-community-modules>LibXML:auth<zef:dwarring>LibXML::Writer:auth<zef:dwarring>Compress::ZlibJSON::Fast:auth<zef:timo>YAMLish:auth<zef:leont>DateTime::Format:auth<zef:raku-community-modules>Syndicate:auth<zef:sasha>

Test Dependencies

Provides

  • Sitemap
  • Sitemap::Builder
  • Sitemap::Config
  • Sitemap::Crawler
  • Sitemap::DirScanner
  • Sitemap::Extraction
  • Sitemap::Fetcher
  • Sitemap::Format::Atom
  • Sitemap::Format::FeedBuilder
  • Sitemap::Format::HTML
  • Sitemap::Format::MRSS
  • Sitemap::Format::RSS
  • Sitemap::Format::TXT
  • Sitemap::Grammar::LinkExtract
  • Sitemap::Grammar::RobotsTxt
  • Sitemap::Grammar::SitemapSniff
  • Sitemap::InputParser
  • Sitemap::Item
  • Sitemap::JsonLd
  • Sitemap::Parser
  • Sitemap::SiteTree
  • Sitemap::Subscribable

The Camelia image is copyright 2009 by Larry Wall. "Raku" is a trademark of the Yet Another Society. All rights reserved.

Built with Podlite — the markup and publishing tools behind this site.