Skip to content

About

Automated B2B lead generation pipeline. Scrapes → LLM-parses → validates → categorizes → stores company data. Zero API costs, local-first, production-grade.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Information Broker Engine

License Python

Automated B2B lead generation and categorization through web scraping and Ollama-powered LLM parsing.

The Problem

Manual B2B lead generation is painfully slow:

  • Time-Consuming: Hours spent scraping company websites manually
  • Inconsistent: Different people extract different information
  • Error-Prone: Human mistakes in data entry and categorization
  • Unscalable: Processing 10-100+ websites requires automation
  • Hallucinations: Unstructured LLM outputs need constant validation

Traditional SaaS lead databases are expensive ($500-1000+/month) and limited to their data sources.

The Solution

Information Broker Engine is an end-to-end B2B lead generation pipeline:

Company Websites → Scrape → Parse (Ollama) → Validate → Categorize → Store

Key Features

✅ Automated Scraping - Extract data from company websites
✅ LLM-Powered Parsing - Mistral via Ollama categorizes and qualifies leads
✅ Structured Output - JSON leads ready for automation
✅ Local Processing - Zero external API calls, complete data privacy
✅ Scalable - Process hundreds of websites efficiently

Results

  • 100% Parsing Success - Structured schemas ensure valid output
  • Deterministic Leads - Consistent tier assignment (hot/warm/cold)
  • Zero API Costs - Local Ollama, no subscription services
  • Production Quality - Validation + cleaning built-in

Architecture

graph LR
    A["Company URLs"] -->|Scrape| B["Raw HTML Content"]
    B -->|Parse| C["Ollama LLM<br/>(Mistral)"]
    C -->|Structure| D["Extracted Data<br/>(JSON)"]
    D -->|Validate| E["Validation Layer<br/>(Schema Check)"]
    E -->|Clean| F["Data Cleaner<br/>(Normalize)"]
    F -->|Store| G["Local Storage<br/>(JSON Files)"]
    G -->|Query| H["Leads Database<br/>(By Tier)"]
    
    style A fill:#e8f4f8
    style G fill:#ccffcc
    style H fill:#ccffcc
Loading

Pipeline Steps

  1. Scrape - Extract HTML content from company website
  2. Parse - Use Ollama/Mistral to intelligently extract company info
  3. Categorize - Classify industry and determine lead tier
  4. Qualify - Score lead and recommend next action
  5. Validate - Check data against schema
  6. Clean - Normalize and standardize fields
  7. Store - Save to local database

Quick Start

Prerequisites

  • Python 3.8+
  • Ollama running locally on http://localhost:11434
  • Mistral model: ollama pull mistral
  • BeautifulSoup4 for web scraping

Installation

git clone https://github.com/buubear14/information-broker-engine.git
cd information-broker-engine
pip install -r requirements.txt

First Example

python examples/example_full_pipeline.py

Expected output:

[00:00] Processing: https://techcorp.example.com
  Step 1: Scraping website... ✓
  Step 2: Parsing with LLM... ✓
  Step 3: Extracting lead qualification... ✓
  Lead score: 87/100 (HOT)

[00:15] Processing: https://dataflow.example.com
  ...

PIPELINE REPORT
├─ Total Leads: 2
├─ Hot Leads: 1
├─ Warm Leads: 1
└─ Average Score: 78

Module Reference

1. WebScraper (src/scraper.py)

Extracts data from company websites.

from src.scraper import WebScraper

scraper = WebScraper()
data = scraper.scrape_website("https://techcorp.com")

print(data['title'])           # Page title
print(data['contact_info'])    # Emails/phones found
print(data['text_content'])    # Main text (first 500 chars)

2. LLMParser (src/parser.py)

Uses Ollama/Mistral to intelligently parse scraped data.

from src.parser import LLMParser

parser = LLMParser()

# Extract structured company data
company_data = parser.parse_company_data(scraped_data)

# Qualify as lead
lead_data = parser.extract_leads(company_data)
print(lead_data['lead_score'])  # 0-100
print(lead_data['tier'])        # hot/warm/cold

3. DataValidator (src/validator.py)

Validates extracted data against schema.

from src.validator import DataValidator

is_valid, errors = DataValidator.validate_lead(lead_data)
if is_valid:
    print("✓ Data is valid")
else:
    print(f"✗ Errors: {errors}")

4. DataCleaner (src/validator.py)

Normalizes and standardizes data.

from src.validator import DataCleaner

cleaned = DataCleaner.clean_lead(raw_lead_data)
# Normalizes strings, validates formats, fixes typos

5. LeadsDatabase (src/database.py)

Stores and retrieves leads.

from src.database import LeadsDatabase

db = LeadsDatabase()
db.add_lead(lead_data)

# Retrieve by tier
hot_leads = db.get_hot_leads()
warm_leads = db.get_warm_leads()

# Export
all_leads = db.get_all()
db.export_to_csv("leads.csv")

6. Pipeline (src/pipeline.py)

End-to-end orchestration.

from src.pipeline import InformationBrokerPipeline

pipeline = InformationBrokerPipeline()
results = pipeline.process_urls([
    "https://company1.com",
    "https://company2.com"
])

report = pipeline.get_report()

Data Schema

Extracted leads conform to this schema (see schemas/lead_schema.json):

{
  "company_name": "TechFlow Inc",
  "website": "https://techflow.com",
  "industry": "SaaS",
  "company_size": "medium",
  "technologies": ["Python", "TensorFlow", "AWS"],
  "contact_email": "sales@techflow.com",
  "lead_score": 87,
  "tier": "hot",
  "next_action": "contact immediately",
  "confidence": 92,
  "data_quality": "high"
}

Field Descriptions

Field Type Values Description
lead_score int 0-100 Lead qualification score
tier string hot/warm/cold Lead tier (hot=contact, warm=nurture, cold=pass)
industry string [Enum] Categorized industry
company_size string startup/small/medium/enterprise Estimated size
confidence int 0-100 Extraction confidence
data_quality string high/medium/low Quality of scraped data

Lead Scoring

Leads are automatically scored and tiered:

Tier Score Action
Hot 75-100 Contact immediately, schedule demo
Warm 50-74 Add to nurture sequence, soft touch
Cold 0-49 Research further, or pass

Scoring factors (determined by LLM):

  • Company size (larger = higher score)
  • Technology match
  • Industry relevance
  • Contact availability
  • Content quality/engagement

Examples

Example 1: Process Single Website

from src.pipeline import InformationBrokerPipeline

pipeline = InformationBrokerPipeline()
lead = pipeline.process_url("https://techcorp.com")
print(f"Score: {lead['lead_score']}/100")

Example 2: Batch Processing

urls = [
    "https://company1.com",
    "https://company2.com",
    "https://company3.com"
]

results = pipeline.process_urls(urls)
print(f"Processed {len(results)} leads")

Example 3: Export to CSV

from src.database import LeadsDatabase

db = LeadsDatabase()
db.export_to_csv("qualified_leads.csv")
# Generates CSV with all leads

Example 4: Query Hot Leads

from src.database import LeadsDatabase

db = LeadsDatabase()
hot_leads = db.get_hot_leads()

for lead in hot_leads:
    print(f"Contact {lead['company_name']} at {lead['contact_email']}")

Performance

Processing times depend on website complexity and Ollama responsiveness:

Step Time
Scrape 1-3s
LLM Parse 5-10s
Validate + Clean 0.5s
Store 0.1s
Total per URL 7-15s

For 100 websites: ~15-20 minutes (can be parallelized for faster results).

Configuration

Customize behavior via environment or code:

from src.pipeline import InformationBrokerPipeline

# Custom storage directory
pipeline = InformationBrokerPipeline(storage_dir="/data/leads")

# Access components
scraper = pipeline.scraper
parser = pipeline.parser
database = pipeline.database

Integration Patterns

Pattern 1: Webhook Handler

from src.pipeline import InformationBrokerPipeline
from flask import Flask, request

app = Flask(__name__)
pipeline = InformationBrokerPipeline()

@app.route('/process_lead', methods=['POST'])
def process_lead():
    url = request.json.get('url')
    lead = pipeline.process_url(url)
    return lead

Pattern 2: Scheduled Job

from src.pipeline import InformationBrokerPipeline
from apscheduler.schedulers.background import BackgroundScheduler

pipeline = InformationBrokerPipeline()
scheduler = BackgroundScheduler()

def process_queue():
    urls = read_urls_from_queue()
    pipeline.process_urls(urls)

scheduler.add_job(process_queue, 'interval', minutes=60)
scheduler.start()

Pattern 3: CLI Tool

from src.pipeline import InformationBrokerPipeline
import click
import json

pipeline = InformationBrokerPipeline()

@click.command()
@click.argument('urls', nargs=-1)
def process(urls):
    results = pipeline.process_urls(list(urls))
    for result in results:
        click.echo(json.dumps(result, indent=2))

if __name__ == '__main__':
    process()

Limitations & Considerations

  1. Scraping Compliance: Respect robots.txt and site ToS
  2. Content Variation: Quality depends on website structure
  3. JavaScript: Static scraping only (no JS rendering)
  4. Rate Limiting: Consider delays between requests
  5. Ollama Required: Must run Ollama locally

Troubleshooting

"Failed to scrape website"

  • Verify URL is accessible
  • Check network connectivity
  • Increase timeout if needed

"Could not parse JSON from response"

  • Ollama may be overloaded
  • Try again or reduce batch size
  • Check Ollama logs

"Validation errors"

  • Data may be incomplete/malformed
  • Check data_quality field
  • Review scraped content

Next Steps

  1. Run Example 1 to see end-to-end flow
  2. Customize for your target websites
  3. Adjust lead scoring if needed
  4. Integrate with your CRM/automation
  5. Deploy for production use

Tech Stack

  • Web Scraping: BeautifulSoup4, Requests
  • LLM: Ollama (Mistral)
  • Storage: Local JSON files
  • Language: Python 3.8+
  • Export: CSV support

Contributing

Areas for enhancement:

  • Async/parallel processing
  • Additional data extractors (social profiles, funding info)
  • Cloud storage support
  • Advanced filtering/querying
  • Integration with CRM systems

License

MIT License - See LICENSE file

Author

Adriaan du Randt - Prompt Engineering & Automation Specialist

Related Projects

  • Structural Prompting Framework: Deterministic prompt engineering
  • Aura Local Agent: Local AI orchestration framework
  • Agent-47: Modular Gemini CLI automation

Transform unstructured web data into qualified B2B leads. Start with Example 1.

About

Automated B2B lead generation pipeline. Scrapes → LLM-parses → validates → categorizes → stores company data. Zero API costs, local-first, production-grade.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages