Skip to content

Repository files navigation

PIB Fact Check Tweet Extraction Pipeline

A multi-stage data engineering pipeline for building a comprehensive, clean dataset of tweets from @PIBFactCheck — the Indian government's official fact-checking handle — using the Twitter/X API v2, the vxTwitter API, and the Wayback Machine CDX API.

Context: This pipeline was developed for a research project on financial misinformation in Indian social media. The goal was to extract every available tweet from PIBFactCheck, recover missing media URLs, patch historical gaps, and resolve truncated retweet text — producing a single clean, deduplicated JSON dataset.


Table of Contents


Overview

The Twitter/X API v2 (Free/Basic tier) imposes a hard cap on how far back you can paginate a user's timeline. This pipeline works around that constraint by using three complementary data sources:

Source Role
Twitter/X API v2 Primary extraction: recent tweets, full text, media metadata
vxTwitter API Free media URL recovery for API gaps; historical tweet content retrieval
Wayback Machine CDX API Index of archived tweet URLs, revealing tweet IDs that predate API pagination limits

The pipeline is designed to be resumable at every stage — if a run is interrupted by a rate limit or connection error, progress is always checkpointed to disk before the exception propagates.


Pipeline Architecture

                        ┌─────────────────────────────┐
                        │   Twitter/X API v2 (Tweepy)  │
                        │   Paginated timeline fetch    │
                        └──────────────┬──────────────┘
                                       │
                              pib_extracted_tweets.json
                              (~3,152 tweets, some null image_url)
                                       │
                        ┌─────────────▼──────────────┐
                        │   vxTwitter API             │
                        │   Patch null image_url      │
                        └──────────────┬──────────────┘
                                       │
                         pib_extracted_tweets_patched.json
                         (504 image URLs recovered)
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   Wayback Machine CDX API                      │
              │   Discover tweet IDs older than API horizon    │
              └────────────────────────┬──────────────────────┘
                                       │
                              711 older tweet IDs found
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   vxTwitter API                                │
              │   Recover content for historical tweet IDs     │
              └────────────────────────┬──────────────────────┘
                                       │
                         pib_extracted_tweets_extended.json
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   Merge + Deduplicate                          │
              │   patched base + older historical tweets       │
              └────────────────────────┬──────────────────────┘
                                       │
                         pib_final_merged.json
                         (~3,863 tweets, deduplicated)
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   Temporal Validation                          │
              │   Snowflake ID → timestamp decoding            │
              │   Gap analysis, monthly distribution           │
              └────────────────────────┬──────────────────────┘
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   Twitter/X API v2 Lookup (batch)              │
              │   Resolve truncated RT texts (referenced_tweets│
              │   expansion)                                   │
              └────────────────────────┬──────────────────────┘
                                       │
                         pib_final_merged_fulltext.json
                                       │
              ┌────────────────────────▼──────────────────────┐
              │   RT Prefix Re-application                     │
              │   "RT @author: " prepended to recovered text   │
              └────────────────────────┬──────────────────────┘
                                       │
                    ✅ pib_final_merged_fulltext_v2.json
                       (Final clean dataset)

Stage-by-Stage Breakdown

Stage 1 — Quick Validation Probe

File: Cell 1
Purpose: Sanity-check API credentials and validate the output schema before committing to a full run.

Fetches 5 tweets with media expansions, builds a media_map from media_key → URL, and prints a formatted JSON payload. This verifies that the Bearer Token is active and that media resolution is working correctly before the bulk loop begins.


Stage 2 — Full Bulk Extraction via Twitter API v2

File: Cell 2
Output: pib_extracted_tweets.json

The core extraction loop using tweepy.Paginator over client.get_users_tweets. Key design points:

  • Checkpoint-and-resume: If the output file already exists and is valid JSON, the loop picks up the until_id from the oldest tweet already saved, continuing backwards in time rather than re-fetching from the top.
  • Rate limit handling: wait_on_rate_limit=True is passed to the Tweepy client, which automatically backs off when the API returns a 429. A 1.5-second politeness sleep is added per batch on top of this.
  • Partial error detection: The API can silently return a 200 with an errors array alongside partial data. The loop explicitly checks response.errors and logs any such warnings without aborting — an important correctness safeguard often omitted in naive implementations.
  • Media resolution: For each batch, a media_map is built from the includes.media expansion and used to resolve image_url for each tweet. Videos resolve to preview_image_url; photos and GIFs resolve to url.
  • Per-batch disk writes: The JSON file is written after every page, not just at the end. This ensures that a mid-run network failure or rate-limit exception loses at most one batch (~100 tweets) of progress.

Each record written:

{
  "tweet_text": "...",
  "image_url": "https://pbs.twimg.com/..." or null,
  "tweet_url": "https://x.com/PIBFactCheck/status/..."
}

Stage 3 — Media URL Patching via vxTwitter

File: Cell 3
Input: pib_extracted_tweets.json
Output: pib_extracted_tweets_patched.json

The Twitter/X API v2 returns null for image_url in some cases where media is present (e.g., certain video formats, some legacy attachment types). This stage iterates over every tweet with a null image_url and queries the vxTwitter API (api.vxtwitter.com) to recover it from the mediaURLs field.

  • Skips tweets that already have a resolved URL.
  • 1.5-second sleep between requests to avoid IP throttling on the free API.
  • Saves the fully patched file at the end of the loop.
  • Result: 504 previously null image URLs were recovered at zero API credit cost.

Stage 4 — Historical Gap Discovery via Wayback Machine CDX API

File: Cell 4
Purpose: Identify tweet IDs that exist in the Internet Archive but are older than what the Twitter API returned.

Queries the Wayback Machine CDX API for all archived snapshots of URLs matching twitter.com/PIBFactCheck/status/*. Extracts tweet IDs from the archived URLs via regex, then diffs against the already-collected set to find IDs that are:

  1. Not yet in the local dataset.
  2. Older (numerically smaller Snowflake ID) than the oldest tweet already collected.

Result: 711 such IDs were found, representing a historical gap the API could not reach.


Stage 5 — Historical Recovery via vxTwitter

File: Cell 5
Input: Older tweet IDs from Stage 4
Output: pib_extracted_tweets_extended.json

Iterates over the 711 discovered IDs (sorted newest-first within the older range) and fetches each one via the vxTwitter API. Media extraction uses a priority fallback chain:

  1. First media_extended entry with type == "image".
  2. Thumbnail URL from media_extended if no direct image found.
  3. First entry from mediaURLs as a last resort.

Progress is checkpointed to disk every 25 tweets. Failed IDs are collected and reported at the end. The final file is the original extracted dataset plus all newly recovered older tweets appended.


Stage 6 — Dataset Merge & Deduplication

File: Cell 6
Inputs: pib_extracted_tweets_patched.json, pib_extracted_tweets_extended.json
Output: pib_final_merged.json

  • Loads the patched base (which has the 504 corrected image URLs from Stage 3).
  • Identifies the 711 new older tweets from the extended file by diffing on tweet_url.
  • Appends only the genuinely new records — preserving the image fixes from Stage 3 rather than overwriting them with potentially worse data from Stage 5.
  • Sanity-checks: total count, unique count, and null image count are printed.

Final count: ~3,863 tweets, fully deduplicated.


Stage 7 — Temporal Validation & Gap Analysis

Files: Cells 7, 8, 9, 10, 11

Decodes Twitter Snowflake IDs to UTC timestamps using the standard formula:

timestamp_ms = (int(snowflake_id) >> 22) + 1288834974657  # Twitter epoch

Runs three analyses:

  1. Date range: Prints earliest and latest tweet dates and the total span covered.
  2. Gap detection: Flags intervals between consecutive tweets greater than 3 days, which may indicate scraping failures or account inactivity periods. Identified a known gap in May–June 2021.
  3. Monthly distribution: Prints a month-by-month tweet count, useful for spotting systematic under-collection in specific periods.
  4. Boundary validation: Identifies the approximate Snowflake ID at the join point between the API-extracted tweets and the Wayback-recovered tweets, and verifies the date aligns with expectations.

Stage 8 — Truncated Retweet Resolution via API v2 Lookup

File: Cell 12
Input: pib_final_merged.json
Output: pib_final_merged_fulltext.json

Twitter's API historically truncated retweet text to 140 characters, appending …. This stage identifies all tweets matching the pattern:

RT @<handle>: <text ending with …>

For each such tweet, it uses client.get_tweets (batch lookups of up to 100 IDs per call) with the referenced_tweets.id expansion. The referenced_tweets array in the response points to the original tweet being retweeted; the includes.tweets expansion returns its full, untruncated text.

The full_text_map is then merged back into the dataset, replacing truncated text with full text where recovered.


Stage 9 — RT Prefix Re-application & Final Cleaning

File: Cell 13
Input: pib_final_merged_fulltext.json
Output: pib_final_merged_fulltext_v2.json

The Stage 8 lookup replaces tweet text with the original tweet's full text — which strips the RT @handle: prefix that was on PIBFactCheck's retweet. This stage restores it:

  • For each tweet where the old text matched ^RT @(\w+): and ended with …, the recovered text is re-prefixed with RT @{author}: .
  • Runs deduplication and truncation-remaining checks as a final sanity pass.

This is the final output file.


Output Schema

Every record in the final dataset conforms to:

{
  "tweet_text": "Full tweet text, including any RT prefix.",
  "image_url": "https://pbs.twimg.com/media/..." ,
  "tweet_url": "https://x.com/PIBFactCheck/status/1234567890123456789"
}
Field Type Notes
tweet_text string Full text; truncated RTs resolved where possible
image_url string | null Primary media URL; null if tweet had no media or media could not be recovered
tweet_url string Canonical X.com URL; the trailing ID segment is a Twitter Snowflake ID

Output Files

File Description
pib_extracted_tweets.json Raw API extraction (~3,152 tweets, some null image_url)
pib_extracted_tweets_patched.json After vxTwitter media patching (504 URLs recovered)
pib_extracted_tweets_extended.json After Wayback + vxTwitter historical recovery (+711 tweets)
pib_final_merged.json Merged, deduplicated dataset (~3,863 tweets)
pib_final_merged_fulltext.json After truncated RT text resolution
pib_final_merged_fulltext_v2.json Final clean dataset — RT prefixes restored, all fixes applied

Prerequisites

Python 3.9+

Install dependencies:

pip install tweepy requests

API Access Required:

API Tier Needed Notes
Twitter/X API v2 Free or Basic Bearer Token only; no OAuth 1.0a required
vxTwitter API None (free, public) No authentication; rate-limited by IP
Wayback Machine CDX API None (free, public) No authentication

Setup & Usage

  1. Clone the repository and open the notebook:

    git clone <repo-url>
    cd <repo-dir>
    jupyter notebook TwitterwithAPI-2.ipynb
  2. Set your Bearer Token in Cells 1, 2, and 12 (the three cells that call the Twitter API):

    BEARER_TOKEN = "your_token_here"

    ⚠️ Do not commit your Bearer Token to version control. Use environment variables or a secrets manager in production.

  3. Run the cells in order. Each stage reads from the output of the previous one. The pipeline is designed so that if any cell is interrupted, re-running it will resume from the last checkpoint rather than starting over.

  4. Adjust the target account by changing the username variable (default: "PIBFactCheck") if you want to run this against a different handle.


Design Decisions & Known Limitations

Why vxTwitter instead of the official API for media patching?
The Twitter/X API v2 Free tier has a monthly tweet cap. Using vxTwitter for image recovery (Stage 3) and historical content retrieval (Stage 5) avoids burning through that quota on non-primary operations. vxTwitter is a read-only proxy with no auth requirement.

Why not just use the Wayback Machine for everything?
Wayback Machine snapshots are inconsistent — not every tweet was archived, and the quality of the HTML snapshot varies. The CDX API is used only for ID discovery, not content extraction. Actual content comes from vxTwitter, which serves structured JSON.

Known data gap (May–June 2021):
Stage 7 analysis revealed a gap in the dataset corresponding to May–June 2021. This gap is confirmed to be genuine (PIBFactCheck had reduced posting activity during this period) and is not a pipeline artifact. The CDX query with limit: 5000 was insufficient to bridge it — a higher limit or a targeted CDX query for that date range could recover additional IDs.

Retweet resolution is partial:
Stage 8 can only resolve truncated RTs for tweets whose referenced tweet IDs are still accessible via the API. Deleted or suspended source tweets will return errors and remain truncated in the final dataset.

Snowflake ID decoding assumes UTC:
All timestamp decoding uses the local system timezone via datetime.fromtimestamp. If running in a non-UTC environment, convert explicitly using datetime.utcfromtimestamp for consistency.


API Cost Summary

Source Cost
Twitter/X API v2 Counted against monthly tweet cap (Free/Basic tier)
vxTwitter API Free (public, IP rate-limited)
Wayback Machine CDX API Free (public)

License

This pipeline is released for research purposes. Usage of the Twitter/X API is subject to the X Developer Agreement and Policy. Data collected via this pipeline should not be redistributed in bulk without compliance with X's data redistribution terms.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages