README-work
WWW::YouTube
Raku package for getting metadata and transcripts of YouTube videos.
The Raku implementation closely follows the Wolfram Language function YouTubeTranscript, [AAf1].
Installation
From Zef ecosystem:
zef install WWW::YouTubeFrom GitHub:
zef install https://github.com/antononcube/Raku-WWW-YouTube.gitUsage
youtube-metadata($id)
Get the metadata of the YouTube video with identifier
$id.
youtube-playlist($id)
Get the video identifiers of the YouTube playlist with identifier
$id.
youtube-transcript($id)
Get the transcript of the YouTube video with identifier
$id.
Details
All three subs,
youtube-metadata,youtube-playlist, andyoutube-transript, work with strings that are identifiers or (full) URLs.youtube-metdataextracts the metadata associated with a YouTube video identifier.Returns a record (hashmap) with keys
<channel-title description publish-date title view-count>.
youtube-playlistextracts the video identifiers of a given YouTube playlist identifier.Currently, gives only the first 100 videos.
youtube-transcriptextracts the captions of the video, if they exist.The transcript can be returned as plain text, array of hashmaps, JSON string.
The YouTube Data API has usage quotas.
Not all YouTube videos have automatic or manual captions. If no captions are available, the function returns a message indicating this.
youtube-transcriptprocesses "captionTracks" of the YouTube Data API, which is a field of YouTube's video metadata.The field "captionTracks" is an array of objects, where each object represents a single caption track (e.g., for a specific language or type).
From "captionTracks" the "baseURL" string is extracted, which is the URL to fetch the caption content.
youtube-transcripthas an option:$methodwhich specifies should the transcript be retrieved using either:Delegation to
youtube_transcript_apiCLI provided by the Python package "youtube-transcript-api", [JDp1]When
:$methodcan be given the values "python", "cli", "python-cli".
Raku's "HTTP::UserAgent"
When
:$methodis given the values "raku", "http", "raku-http".
Examples
Metadata
Get the metadata associated with a YouTube video identifier:
use WWW::YouTube;
use Data::Translators;
youtube-metadata('S_3e7liz4KM')
==> to-html(align => 'left')Transcripts
my $transcript = youtube-transcript('ewU83vHwN8Y');
say $transcript.chars;
say $transcript.substr(^300);Summarize using a Large Language Model (LLM):
use LLM::Functions;
use LLM::Prompts;
llm-synthesize(llm-prompt('Summarize')($transcript), e => 'ChatGPT')Get the transcript as a dataset:
my @t = youtube-transcript('S_3e7liz4KM', format => 'dataset');
@t.head(10) ==> to-html(field-names => <time duration content>, align => 'left')Playlists
youtube-playlist('PLke9UbqjOSOiMnn8kNg6pb3TFWDsqjNTN')CLI
The package provides Command Line Interface (CLI) scripts. Here are their usage messages:
youtube-metadata --helpyoutube-playlist --helpyoutube-transcript --helpTODO
TODO Implementation
DONE Get transcript for a video identifier
DONE Video metadata retrieval
TODO Video identifiers for a playlist
DONE For playlists with ⤠100 videos
TODO Large playlists
Delegation to the Python package "youtube-transcript-api"
TODO Different transcript output formats (for the Raku-HTTP method)
DONE Text
DONE Dataset (array of hashmap records)
DONE JSON
TODO WebVTT
TODO SRT
Implement versions of the subs using a YouTube API key
TODO Documentation
DONE Basic usage
TODO Transcripts retrieval for a playlist
References
[AAf1] Anton Antonov, YouTubeTranscript, (2025), Wolfram Function Repository.
[JDp1] Jonas Depoix, youtube-transcript-api Python package, (2018-2025), GitHub/jdepoix. (At PyPI.org.)