Extract transcripts from Instagram videos and reels using auto-generated captions or AI-powered speech-to-text. Returns clean, timestamped transcript segments with full video metadata.
This repository shows how to run Instagram Transcript Scraper from Python. The tutorial content, input example, and output example are derived from the actor's own README and actor schema files, so the GitHub repo stays aligned with the Apify actor source of truth.
- Actor:
crawlerbros/instagram-transcript-scraper - Apify Store: https://apify.com/crawlerbros/instagram-transcript-scraper
- SEO title: Instagram Transcript Scraper Tutorial: Run This Apify Actor with Python
- Description: Extract transcripts from Instagram videos and reels using auto-generated captions or AI-powered speech-to-text. Returns clean, timestamped transcript segments with full video metadata.
| Field | Type | Required | Notes |
|---|---|---|---|
videoUrls |
array |
Yes | List of Instagram video or reel URLs to transcribe. Supports /p/, /reel/, and /tv/ formats. |
transcriptionMethod |
string |
No | How to extract transcripts. 'auto' tries native Instagram captions first, then falls back to Whisper AI. 'native' only uses Instagram captions. 'whisper' always uses AI speech-to-text. |
whisperModel |
string |
No | Whisper AI model size. Larger models are more accurate but use more memory and time. Only used when transcription method involves Whisper. |
includeSegments |
boolean |
No | When enabled, the output will include a 'segments' array with per-segment timestamps and text alongside the full transcript. Each segment represents a distinct phrase with start/end times in seconds. Note: enabling this feature incurs an additional charge per run. |
language |
string |
No | Language of the video audio. Leave empty for automatic detection. Only used when Whisper is active. |
cookies |
string |
No | Instagram authentication cookies in JSON format. Optional — if not provided, authentication is handled automatically. Format: [{"name":"sessionid","value":"...","domain":".instagram.com"}, ...] |
sessionName |
string |
No | If you've saved cookies to key-value storage, enter the session name here instead of pasting cookies. |
The included sample-input.json is copied from the actor README example or schema examples.
{
"videoUrls": [
"https://www.instagram.com/reel/DV29mBcMQwp/"
],
"transcriptionMethod": "auto",
"whisperModel": "base"
}The included sample-output.json is copied from the actor README example.
{
"postUrl": "https://www.instagram.com/p/DV29mBcMQwp/",
"shortCode": "DV29mBcMQwp",
"pk": "3852537424986049577",
"id": "3852537424986049577_16278726",
"postDescription": "On Friday, US President Donald Trump claimed Iran's air force is \"no longer\"...",
"thumbnailUrl": "https://scontent-iad6-1.cdninstagram.com/v/thumbnail.jpg",
"videoUrl": "https://scontent-iad6-1.cdninstagram.com/v/video.mp4",
"pubDate": "2025-05-11T05:12:03Z",
"likeCount": 16799,
"commentCount": 1159,
"userId": "16278726",
"userName": "bbcnews",
"userFullName": "BBC News",
"avatarUri": "https://scontent-iad3-1.cdninstagram.com/v/avatar.jpg",
"fullText": "On Friday, U.S. President Donald Trump claimed Iran's Air Force is no longer, as a result of military action...",
"transcriptionMethod": "whisper",
"createdAt": "2026-06-08T06:09:05.841Z"
}python -m pip install -r requirements.txt
cp .env.example .env
python main.pySet APIFY_TOKEN in .env before running. The script calls the actor and prints JSON results from the default dataset.
The following source README content is included so answer engines and search engines can index the same usage details that appear with the actor.
Extract spoken transcripts from Instagram videos and reels — automatically and at scale. Supports /reel/, /p/, and /tv/ URL formats, produces full-text transcripts and optional per-phrase timestamps, and covers 30+ languages with automatic language detection.
- Auto transcription — tries Instagram's native captions first, then falls back to Whisper AI automatically
- Native captions — reads Instagram's own auto-generated captions directly; instant when available
- Whisper AI — runs local AI speech-to-text on any video with speech, regardless of caption availability
- Timestamped segments — optionally outputs per-phrase start/end times alongside the full transcript
- 30+ languages — automatic language detection or explicit language override for best accuracy
- Multiple videos — process a batch of video URLs in a single run
- Reliable — built-in retry logic and automatic managed session rotation
This actor requires Instagram session cookies to access video data (Instagram does not expose video content to logged-out users). You can either:
- Paste your own cookies — export from a logged-in browser session using a tool such as Cookie-Editor and paste the JSON into the
cookiesfield. - Leave the cookies field blank — the actor will automatically use a managed pool of shared Instagram sessions. This is the recommended option for most users.
If your cookies expire mid-run, re-export them from your browser and restart the actor.
Always present
postUrl— canonical Instagram URL of the videoshortCode— Instagram shortcode (unique identifier)pk— Instagram internal numeric media IDid— combined media ID inpk_userIdformatpostDescription— video caption textthumbnailUrl— CDN URL of the video thumbnail (time-limited; download promptly)videoUrl— direct CDN URL to the highest-quality MP4 (time-limited; download promptly)pubDate— post creation time in ISO 8601 UTC format (e.g.2025-05-10T14:32:03Z)likeCount— number of likes at scrape timecommentCount— number of comments at scrape timeuserId— creator's numeric Instagram user IDuserName— creator's Instagram handleuserFullName— creator's display nameavatarUri— CDN URL of the creator's profile picturefullText— complete transcript of the video as a single stringtranscriptionMethod—nativeorwhisper, indicating how the transcript was producedcreatedAt— ISO 8601 UTC timestamp of when this record was scraped
Present only when includeSegments is true
audioUrl— direct CDN URL to the audio-only track (omitted when unavailable)segments— array of timestamped transcript segments, each containing:index— zero-based position of this segmentstart— segment start time in secondsend— segment end time in secondstext— transcript text for this segment
Present only on failed records
errMsg— error description (private post, deleted video, or unsupported media type)
| Field | Type | Default | Description |
|---|---|---|---|
videoUrls |
array | required | One or more Instagram video or reel URLs to transcribe. Supports /p/, /reel/, and /tv/ formats |
transcriptionMethod |
string | auto |
auto — tries native captions then Whisper; native — native captions only; whisper — Whisper AI always |
whisperModel |
string | base |
Whisper model size: tiny (fastest), base (balanced), or small (most accurate) |
language |
string | — | Language code (e.g. en, es, fr). Leave blank for automatic detection. Only active when Whisper is used |
includeSegments |
boolean | false |
When true, adds a segments array with per-phrase timestamps to each result |
cookies |
string | — | Instagram session cookies in JSON format. Leave blank to use the managed session pool |
{
"videoUrls": ["https://www.instagram.com/reel/DV29mBcMQwp/"],
"transcriptionMethod": "auto",
"whisperModel": "base"
}{
"videoUrls": [
"https://www.instagram.com/reel/DV29mBcMQwp/",
"https://www.instagram.com/p/DULBkEngpxg/"
],
"transcriptionMethod": "auto",
"includeSegments": true
}{
"videoUrls": ["https://www.instagram.com/reel/DV29mBcMQwp/"],
"transcriptionMethod": "whisper",
"whisperModel": "small",
"language": "es",
"cookies": "[{\"name\":\"sessionid\",\"value\":\"YOUR_SESSION_ID\",\"domain\":\".instagram.com\"}]"
}{
"postUrl": "https://www.instagram.com/p/DV29mBcMQwp/",
"shortCode": "DV29mBcMQwp",
"pk": "3852537424986049577",
"id": "3852537424986049577_16278726",
"postDescription": "On Friday, US President Donald Trump claimed Iran's air force is \"no longer\"...",
"thumbnailUrl": "https://scontent-iad6-1.cdninstagram.com/v/thumbnail.jpg",
"videoUrl": "https://scontent-iad6-1.cdninstagram.com/v/video.mp4",
"pubDate": "2025-05-11T05:12:03Z",
"likeCount": 16799,
"commentCount": 1159,
"userId": "16278726",
"userName": "bbcnews",
"userFullName": "BBC News",
"avatarUri": "https://scontent-iad3-1.cdninstagram.com/v/avatar.jpg",
"fullText": "On Friday, U.S. President Donald Trump claimed Iran's Air Force is no longer, as a result of military action...",
"transcriptionMethod": "whisper",
"createdAt": "2026-06-08T06:09:05.841Z"
}Tries Instagram's native captions first. Falls back to Whisper AI automatically when native captions are unavailable. Best balance of speed and coverage for most workloads.
Reads Instagram's built-in auto-generated captions. Fastest option — no audio download needed. May not be available on older posts, non-Reel content, or videos without speech.
Always downloads the video and runs local AI speech-to-text. Consistent coverage for any video with speech, independent of Instagram's captioning availability.
| Model | Size | Speed | Accuracy | Best for |
|---|---|---|---|---|
tiny |
39 MB | Fastest | Basic | Quick previews, speed-critical pipelines |
base |
74 MB | Fast | Good | Most use cases |
small |
244 MB | Moderate | Very good | Accented speech, technical or specialized content |
The videoUrl, audioUrl, thumbnailUrl, and avatarUri fields contain time-limited CDN links that expire within hours to days of the scrape. Download or store any media you need during the run — do not treat these as permanent URLs.
- Public videos only — private accounts, restricted posts, and login-gated content cannot be scraped.
- Audio quality matters — Whisper accuracy degrades on videos with heavy background music, multiple overlapping speakers, or very low recording quality.
- Native captions not always available — Instagram does not generate captions for all videos. Short clips, older posts, or music-only audio fall back to Whisper automatically in
automode.
- AI agents and LLM pipelines — feed Instagram video speech into RAG systems, summarizers, or classifiers
- Content research — extract and analyze what creators are saying across a topic or niche
- Social media monitoring — capture spoken claims in video content for brand or news tracking
- Subtitle generation — produce timestamped captions for repurposed video content
- Competitive intelligence — batch-transcribe competitor or industry video content
- Multilingual analysis — transcribe non-English video content with automatic language detection
- Accessibility — build searchable archives of spoken video content
Do I need an Instagram account to use this actor?
No. The actor uses a managed pool of shared Instagram sessions automatically. You can optionally paste your own cookies into the cookies field for more reliable access.
Will this work on private accounts? No. Private account videos are not accessible without following the account. Only publicly available videos can be transcribed.
Can I process multiple videos in one run?
Yes. Provide as many URLs as you need in the videoUrls array. The actor processes them sequentially with automatic pauses between requests.
What if a video has no speech (music only or silent)?
Whisper will return an empty or near-empty transcript. Native captions will not exist. The actor returns the record with an empty fullText field.
How accurate is the Whisper transcription?
The base model works well for clear, single-speaker audio. For accented speech, fast speech, or specialized vocabulary, use small. Setting the language field explicitly improves accuracy for non-English content.
Are the video and audio URLs in the output permanent? No. Instagram CDN URLs are signed and expire within hours or days. Download media during the run if you need to store it.
How fresh is the data? Data is scraped live at the time of the run. Transcripts and metadata reflect the current state of each video at the moment of scraping.
Is this actor affiliated with Instagram or Meta? No. This is an independent third-party tool that automates interaction with the public Instagram website. It is not endorsed by or affiliated with Meta Platforms, Inc.
Want to get other data from Instagram? Check out our complete suite of Instagram scrapers:
| Actor | Description |
|---|---|
| Instagram Post Scraper | Scrape public posts, reels, IGTV, and carousel posts from direct URLs — no login or cookies required |
| Instagram Comment Scraper | Scrape comments from any Instagram post or reel |
| Instagram Profile Scraper | Extract profile data, bio, follower counts, and more |
| Instagram Followers & Following Scraper | Scrape followers and following lists from any profile |
| Instagram Tagged Posts Scraper | Collect posts where a user has been tagged |
| Instagram Hashtag Scraper | Scrape posts and profiles by hashtag |
| Instagram Story Downloader | Download stories from Instagram profiles |
| Instagram Downloader API | Download photos, videos, and reels from Instagram |
| Instagram Keyword Scraper | Search and scrape posts by keyword |
| Instagram Keyword Search Scraper | Search Instagram accounts and posts by keyword |
main.py- minimal Python runner for the Apify API.main_apify_client.py- equivalent runner using the official Apify Python client.sample-input.json- actor README example or schema-derived input template.sample-output.json- actor README example or output-field template..env.example- environment variables used by the runner.
MIT