This project now includes shell scripts to automate everything:
| Script | Purpose | When to Use |
|---|---|---|
run_pipeline.sh |
Complete pipeline execution | Main script - does everything |
start_docker.sh |
Start Docker container | Before running anything |
start_services.sh |
Start Hadoop services | First time setup inside container |
cd "/home/ryukr2/Projects/ClickSteam analysis"
chmod +x run_pipeline.sh start_docker.sh start_services.sh./start_docker.sh amd # Use 'arm' for M1/M2 Mac
# This automatically:
# ✓ Creates/starts container
# ✓ Mounts project folder
# ✓ Opens interactive shell inside container# Inside container:
./start_services.sh
# This automatically:
# ✓ Formats NameNode
# ✓ Starts NameNode, DataNode
# ✓ Starts ResourceManager, NodeManager
# ✓ Starts Hive MetaStore
# ✓ Shows all running services# Inside container, run once:
hdfs dfs -mkdir -p /user/root/clickstream/{raw,processed}# Inside container:
./run_pipeline.sh
# This automatically:
# ✓ Generates 100 sample logs
# ✓ Uploads to HDFS
# ✓ Runs Pig cleaning (removes 11% bad data)
# ✓ Creates Hive table
# ✓ Runs 8 analytics queries
# ✓ Displays results
# Total time: ~2 minutes./run_pipeline.sh # Run complete pipeline
./run_pipeline.sh --generate # Only generate logs
./run_pipeline.sh --upload # Only upload to HDFS
./run_pipeline.sh --clean # Only clean data with Pig
./run_pipeline.sh --analyze # Only run analytics
./run_pipeline.sh --help # Show all optionsProcess Flow:
Generate 100 logs
↓
Upload to HDFS raw directory
↓
Delete old processed data
↓
Run Pig ETL (MapReduce job)
↓
Create Hive table schema
↓
Run 8 analytics queries
↓
Display results
Features:
- ✅ Colored output (helps read progress)
- ✅ Error checking at each step
- ✅ Detailed logging
- ✅ Automatic cleanup
- ✅ Can run individual steps
./start_docker.sh amd # Start for AMD/Intel Linux
./start_docker.sh arm # Start for Mac M1/M2What It Does:
- Checks if container exists
- If not, creates new container with:
- Port mappings (9870, 8088, 9864, 9083)
- Volume mount (project folder)
- keeps running in background
- Connects you to interactive shell
Smart Features:
- Reuses existing container if running
- Auto-restarts stopped container
- Handles both AMD and ARM architectures
./start_services.shSteps:
- Format NameNode (creates file system)
- Start NameNode daemon
- Start DataNode daemon
- Start ResourceManager (YARN)
- Start NodeManager
- Start Hive MetaStore
- Verify with
jpscommand - Show web UI URLs
Output Shows:
NameNode listening on port 9870
ResourceManager listening on port 8088
DataNode listening on port 9864
Hive MetaStore listening on port 9083
# Edit crontab
crontab -e
# Add this line:
0 0 * * * cd /home/ryukr2/Projects/ClickSteam\ analysis && docker exec clickstream /clickstream/run_pipeline.sh >> /tmp/pipeline.log 2>&1What This Does:
- Runs pipeline every day at 00:00 (midnight)
- Logs output to
/tmp/pipeline.log - Runs inside Docker container automatically
Log File:
# View logs
tail -f /tmp/pipeline.log
# See last 100 lines
tail -100 /tmp/pipeline.log# Every hour at minute 0
0 * * * * cd /home/ryukr2/Projects/ClickSteam\ analysis && docker exec clickstream /clickstream/run_pipeline.sh >> /tmp/pipeline.log 2>&1# Every 15 minutes
*/15 * * * * cd /home/ryukr2/Projects/ClickSteam\ analysis && docker exec clickstream /clickstream/run_pipeline.sh >> /tmp/pipeline.log 2>&1# 1. Start services
./start_services.sh
# 2. Create directories
hdfs dfs -mkdir -p /user/root/clickstream/{raw,processed}
# 3. Schedule pipeline
crontab -e
# Add: 0 0 * * * cd /path/to/project && docker exec clickstream /clickstream/run_pipeline.sh >> /tmp/pipeline.log 2>&1Every day at 00:00 (midnight):
1. Generate 100 new log entries
2. Upload to HDFS
3. Clean data (remove 11% bad records)
4. Run 8 analytics queries
5. Results appear (could save to file)
Results available for morning review!
Step 1: Copy your logs to project
cp /var/log/apache2/access.log /home/ryukr2/Projects/ClickSteam\ analysis/logs/Step 2: Modify run_pipeline.sh
# Change this line:
# FROM: python3 << 'PYTHON_EOF' (generates fake logs)
# TO: # Skip log generation, use real logs
# Just remove or comment out the generate_logs() callStep 3: Run pipeline
./run_pipeline.shCreate email script:
# Create: email_results.sh
RESULTS="/tmp/pipeline_results.txt"
hive -hiveconf hive.metastore.uris=thrift://localhost:9083 \
-f /clickstream/phase3_analysis/trend_queries.hql > "$RESULTS" 2>&1
# Email results
cat "$RESULTS" | mail -s "Daily Clickstream Analysis" your_email@example.comUpdate crontab:
0 0 * * * cd /path && ./email_results.sh# Make sure it's executable
chmod +x *.sh
# Check syntax errors
bash -n run_pipeline.sh
# Run with debug output
bash -x run_pipeline.sh# Check if running
docker ps | grep clickstream
# View container logs
docker logs clickstream
# Stop container
docker stop clickstream
# Start container
docker start clickstream
# Remove container (and start fresh)
docker rm clickstream
./start_docker.sh amd# Check services
jps
# Check NameNode health
hdfs dfsadmin -report
# View Hive MetaStore log
cat /tmp/metastore.log┌─ Every Day at Midnight ─┐
│ │
└─→ run_pipeline.sh │
├─ Generate logs │
├─ Upload to HDFS │
├─ Clean with Pig │
├─ Run Hive queries │
└─ Display results │
│
Morning:
Results ready
for review!
- Make scripts executable:
chmod +x *.sh - Test
start_docker.sh - Test
start_services.sh - Test
run_pipeline.shmanually - Create HDFS directories:
hdfs dfs -mkdir -p /user/root/clickstream/{raw,processed} - Test full pipeline once:
./run_pipeline.sh - Schedule with crontab:
crontab -e - Verify cron job:
crontab -l - Check logs:
tail -f /tmp/pipeline.log
./run_pipeline.sh
# Output: ~2 minutes, shows all query results./run_pipeline.sh --generate
# Output: 100 logs generated./run_pipeline.sh --analyze
# Output: 8 queries executeddocker rm clickstream # Remove old
./start_docker.sh amd # Create new# View all cron jobs
crontab -l
# View cron log
log show --predicate 'process == "cron"' --last 1h
# View custom pipeline log
tail -f /tmp/pipeline.log| Before Scripts | After Scripts |
|---|---|
| Manual commands Each step typed separately Easy to forget steps Time-consuming |
Automated execution One command does all Consistent, repeatable Saves hours per month |
| No scheduling Must run manually |
Runs automatically daily Results ready every morning Zero manual effort |
| Hard to debug Errors stop process manually |
Error checking at each step Colored output for clarity Logs saved for review |
- Test Scripts: Run each script manually first
- Verify Output: Check that pipeline produces expected results
- Schedule Job: Add to crontab for automation
- Monitor Logs: Review
/tmp/pipeline.logdaily - Scale Up: Gradually increase log size or frequency
Your pipeline is now fully automated ⚡