A multi-stage data engineering pipeline for building a comprehensive, clean dataset of tweets from @PIBFactCheck — the Indian government's official fact-checking handle — using the Twitter/X API v2, the vxTwitter API, and the Wayback Machine CDX API.
Context: This pipeline was developed for a research project on financial misinformation in Indian social media. The goal was to extract every available tweet from PIBFactCheck, recover missing media URLs, patch historical gaps, and resolve truncated retweet text — producing a single clean, deduplicated JSON dataset.
- Overview
- Pipeline Architecture
- Stage-by-Stage Breakdown
- Stage 1 — Quick Validation Probe
- Stage 2 — Full Bulk Extraction via Twitter API v2
- Stage 3 — Media URL Patching via vxTwitter
- Stage 4 — Historical Gap Discovery via Wayback Machine CDX API
- Stage 5 — Historical Recovery via vxTwitter
- Stage 6 — Dataset Merge & Deduplication
- Stage 7 — Temporal Validation & Gap Analysis
- Stage 8 — Truncated Retweet Resolution via API v2 Lookup
- Stage 9 — RT Prefix Re-application & Final Cleaning
- Output Schema
- Output Files
- Prerequisites
- Setup & Usage
- Design Decisions & Known Limitations
- API Cost Summary
The Twitter/X API v2 (Free/Basic tier) imposes a hard cap on how far back you can paginate a user's timeline. This pipeline works around that constraint by using three complementary data sources:
| Source | Role |
|---|---|
| Twitter/X API v2 | Primary extraction: recent tweets, full text, media metadata |
| vxTwitter API | Free media URL recovery for API gaps; historical tweet content retrieval |
| Wayback Machine CDX API | Index of archived tweet URLs, revealing tweet IDs that predate API pagination limits |
The pipeline is designed to be resumable at every stage — if a run is interrupted by a rate limit or connection error, progress is always checkpointed to disk before the exception propagates.
┌─────────────────────────────┐
│ Twitter/X API v2 (Tweepy) │
│ Paginated timeline fetch │
└──────────────┬──────────────┘
│
pib_extracted_tweets.json
(~3,152 tweets, some null image_url)
│
┌─────────────▼──────────────┐
│ vxTwitter API │
│ Patch null image_url │
└──────────────┬──────────────┘
│
pib_extracted_tweets_patched.json
(504 image URLs recovered)
│
┌────────────────────────▼──────────────────────┐
│ Wayback Machine CDX API │
│ Discover tweet IDs older than API horizon │
└────────────────────────┬──────────────────────┘
│
711 older tweet IDs found
│
┌────────────────────────▼──────────────────────┐
│ vxTwitter API │
│ Recover content for historical tweet IDs │
└────────────────────────┬──────────────────────┘
│
pib_extracted_tweets_extended.json
│
┌────────────────────────▼──────────────────────┐
│ Merge + Deduplicate │
│ patched base + older historical tweets │
└────────────────────────┬──────────────────────┘
│
pib_final_merged.json
(~3,863 tweets, deduplicated)
│
┌────────────────────────▼──────────────────────┐
│ Temporal Validation │
│ Snowflake ID → timestamp decoding │
│ Gap analysis, monthly distribution │
└────────────────────────┬──────────────────────┘
│
┌────────────────────────▼──────────────────────┐
│ Twitter/X API v2 Lookup (batch) │
│ Resolve truncated RT texts (referenced_tweets│
│ expansion) │
└────────────────────────┬──────────────────────┘
│
pib_final_merged_fulltext.json
│
┌────────────────────────▼──────────────────────┐
│ RT Prefix Re-application │
│ "RT @author: " prepended to recovered text │
└────────────────────────┬──────────────────────┘
│
✅ pib_final_merged_fulltext_v2.json
(Final clean dataset)
File: Cell 1
Purpose: Sanity-check API credentials and validate the output schema before committing to a full run.
Fetches 5 tweets with media expansions, builds a media_map from media_key → URL, and prints a formatted JSON payload. This verifies that the Bearer Token is active and that media resolution is working correctly before the bulk loop begins.
File: Cell 2
Output: pib_extracted_tweets.json
The core extraction loop using tweepy.Paginator over client.get_users_tweets. Key design points:
- Checkpoint-and-resume: If the output file already exists and is valid JSON, the loop picks up the
until_idfrom the oldest tweet already saved, continuing backwards in time rather than re-fetching from the top. - Rate limit handling:
wait_on_rate_limit=Trueis passed to the Tweepy client, which automatically backs off when the API returns a 429. A 1.5-second politeness sleep is added per batch on top of this. - Partial error detection: The API can silently return a 200 with an
errorsarray alongside partial data. The loop explicitly checksresponse.errorsand logs any such warnings without aborting — an important correctness safeguard often omitted in naive implementations. - Media resolution: For each batch, a
media_mapis built from theincludes.mediaexpansion and used to resolveimage_urlfor each tweet. Videos resolve topreview_image_url; photos and GIFs resolve tourl. - Per-batch disk writes: The JSON file is written after every page, not just at the end. This ensures that a mid-run network failure or rate-limit exception loses at most one batch (~100 tweets) of progress.
Each record written:
{
"tweet_text": "...",
"image_url": "https://pbs.twimg.com/..." or null,
"tweet_url": "https://x.com/PIBFactCheck/status/..."
}File: Cell 3
Input: pib_extracted_tweets.json
Output: pib_extracted_tweets_patched.json
The Twitter/X API v2 returns null for image_url in some cases where media is present (e.g., certain video formats, some legacy attachment types). This stage iterates over every tweet with a null image_url and queries the vxTwitter API (api.vxtwitter.com) to recover it from the mediaURLs field.
- Skips tweets that already have a resolved URL.
- 1.5-second sleep between requests to avoid IP throttling on the free API.
- Saves the fully patched file at the end of the loop.
- Result: 504 previously null image URLs were recovered at zero API credit cost.
File: Cell 4
Purpose: Identify tweet IDs that exist in the Internet Archive but are older than what the Twitter API returned.
Queries the Wayback Machine CDX API for all archived snapshots of URLs matching twitter.com/PIBFactCheck/status/*. Extracts tweet IDs from the archived URLs via regex, then diffs against the already-collected set to find IDs that are:
- Not yet in the local dataset.
- Older (numerically smaller Snowflake ID) than the oldest tweet already collected.
Result: 711 such IDs were found, representing a historical gap the API could not reach.
File: Cell 5
Input: Older tweet IDs from Stage 4
Output: pib_extracted_tweets_extended.json
Iterates over the 711 discovered IDs (sorted newest-first within the older range) and fetches each one via the vxTwitter API. Media extraction uses a priority fallback chain:
- First
media_extendedentry withtype == "image". - Thumbnail URL from
media_extendedif no direct image found. - First entry from
mediaURLsas a last resort.
Progress is checkpointed to disk every 25 tweets. Failed IDs are collected and reported at the end. The final file is the original extracted dataset plus all newly recovered older tweets appended.
File: Cell 6
Inputs: pib_extracted_tweets_patched.json, pib_extracted_tweets_extended.json
Output: pib_final_merged.json
- Loads the patched base (which has the 504 corrected image URLs from Stage 3).
- Identifies the 711 new older tweets from the extended file by diffing on
tweet_url. - Appends only the genuinely new records — preserving the image fixes from Stage 3 rather than overwriting them with potentially worse data from Stage 5.
- Sanity-checks: total count, unique count, and null image count are printed.
Final count: ~3,863 tweets, fully deduplicated.
Files: Cells 7, 8, 9, 10, 11
Decodes Twitter Snowflake IDs to UTC timestamps using the standard formula:
timestamp_ms = (int(snowflake_id) >> 22) + 1288834974657 # Twitter epochRuns three analyses:
- Date range: Prints earliest and latest tweet dates and the total span covered.
- Gap detection: Flags intervals between consecutive tweets greater than 3 days, which may indicate scraping failures or account inactivity periods. Identified a known gap in May–June 2021.
- Monthly distribution: Prints a month-by-month tweet count, useful for spotting systematic under-collection in specific periods.
- Boundary validation: Identifies the approximate Snowflake ID at the join point between the API-extracted tweets and the Wayback-recovered tweets, and verifies the date aligns with expectations.
File: Cell 12
Input: pib_final_merged.json
Output: pib_final_merged_fulltext.json
Twitter's API historically truncated retweet text to 140 characters, appending …. This stage identifies all tweets matching the pattern:
RT @<handle>: <text ending with …>
For each such tweet, it uses client.get_tweets (batch lookups of up to 100 IDs per call) with the referenced_tweets.id expansion. The referenced_tweets array in the response points to the original tweet being retweeted; the includes.tweets expansion returns its full, untruncated text.
The full_text_map is then merged back into the dataset, replacing truncated text with full text where recovered.
File: Cell 13
Input: pib_final_merged_fulltext.json
Output: pib_final_merged_fulltext_v2.json
The Stage 8 lookup replaces tweet text with the original tweet's full text — which strips the RT @handle: prefix that was on PIBFactCheck's retweet. This stage restores it:
- For each tweet where the old text matched
^RT @(\w+):and ended with…, the recovered text is re-prefixed withRT @{author}:. - Runs deduplication and truncation-remaining checks as a final sanity pass.
This is the final output file.
Every record in the final dataset conforms to:
{
"tweet_text": "Full tweet text, including any RT prefix.",
"image_url": "https://pbs.twimg.com/media/..." ,
"tweet_url": "https://x.com/PIBFactCheck/status/1234567890123456789"
}| Field | Type | Notes |
|---|---|---|
tweet_text |
string |
Full text; truncated RTs resolved where possible |
image_url |
string | null |
Primary media URL; null if tweet had no media or media could not be recovered |
tweet_url |
string |
Canonical X.com URL; the trailing ID segment is a Twitter Snowflake ID |
| File | Description |
|---|---|
pib_extracted_tweets.json |
Raw API extraction (~3,152 tweets, some null image_url) |
pib_extracted_tweets_patched.json |
After vxTwitter media patching (504 URLs recovered) |
pib_extracted_tweets_extended.json |
After Wayback + vxTwitter historical recovery (+711 tweets) |
pib_final_merged.json |
Merged, deduplicated dataset (~3,863 tweets) |
pib_final_merged_fulltext.json |
After truncated RT text resolution |
pib_final_merged_fulltext_v2.json |
Final clean dataset — RT prefixes restored, all fixes applied |
Python 3.9+
Install dependencies:
pip install tweepy requestsAPI Access Required:
| API | Tier Needed | Notes |
|---|---|---|
| Twitter/X API v2 | Free or Basic | Bearer Token only; no OAuth 1.0a required |
| vxTwitter API | None (free, public) | No authentication; rate-limited by IP |
| Wayback Machine CDX API | None (free, public) | No authentication |
-
Clone the repository and open the notebook:
git clone <repo-url> cd <repo-dir> jupyter notebook TwitterwithAPI-2.ipynb
-
Set your Bearer Token in Cells 1, 2, and 12 (the three cells that call the Twitter API):
BEARER_TOKEN = "your_token_here"
⚠️ Do not commit your Bearer Token to version control. Use environment variables or a secrets manager in production. -
Run the cells in order. Each stage reads from the output of the previous one. The pipeline is designed so that if any cell is interrupted, re-running it will resume from the last checkpoint rather than starting over.
-
Adjust the target account by changing the
usernamevariable (default:"PIBFactCheck") if you want to run this against a different handle.
Why vxTwitter instead of the official API for media patching?
The Twitter/X API v2 Free tier has a monthly tweet cap. Using vxTwitter for image recovery (Stage 3) and historical content retrieval (Stage 5) avoids burning through that quota on non-primary operations. vxTwitter is a read-only proxy with no auth requirement.
Why not just use the Wayback Machine for everything?
Wayback Machine snapshots are inconsistent — not every tweet was archived, and the quality of the HTML snapshot varies. The CDX API is used only for ID discovery, not content extraction. Actual content comes from vxTwitter, which serves structured JSON.
Known data gap (May–June 2021):
Stage 7 analysis revealed a gap in the dataset corresponding to May–June 2021. This gap is confirmed to be genuine (PIBFactCheck had reduced posting activity during this period) and is not a pipeline artifact. The CDX query with limit: 5000 was insufficient to bridge it — a higher limit or a targeted CDX query for that date range could recover additional IDs.
Retweet resolution is partial:
Stage 8 can only resolve truncated RTs for tweets whose referenced tweet IDs are still accessible via the API. Deleted or suspended source tweets will return errors and remain truncated in the final dataset.
Snowflake ID decoding assumes UTC:
All timestamp decoding uses the local system timezone via datetime.fromtimestamp. If running in a non-UTC environment, convert explicitly using datetime.utcfromtimestamp for consistency.
| Source | Cost |
|---|---|
| Twitter/X API v2 | Counted against monthly tweet cap (Free/Basic tier) |
| vxTwitter API | Free (public, IP rate-limited) |
| Wayback Machine CDX API | Free (public) |
This pipeline is released for research purposes. Usage of the Twitter/X API is subject to the X Developer Agreement and Policy. Data collected via this pipeline should not be redistributed in bulk without compliance with X's data redistribution terms.