Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Server Operations Toolkit

A production-ready collection of Bash automation utilities designed to help DevOps Engineers automate routine operational tasks across Ubuntu Linux servers and AWS EC2 environments.

1.0 Project Overview

A production-ready collection of Bash automation utilities designed to help DevOps Engineers automate routine operational tasks across Ubuntu Linux servers and AWS EC2 environments. The toolkit focuses on infrastructure auditing, service management, storage analysis, network validation, and backup automation through modular, production-oriented shell scripts that can be executed independently or integrated into existing operational workflows.

Each utility addresses a specific aspect of day-to-day infrastructure operations, enabling engineers to perform server health assessments, audit network configurations, monitor critical services, analyse storage utilisation, and manage backup lifecycles using a consistent command-line interface. The scripts are designed to support common operational activities in development, testing, and production environments while remaining lightweight, reusable, and easy to customise for different infrastructure requirements.

2.0 Repository Structure

server-operations-toolkit/
│
├── README.md
├── LICENSE
├── .gitignore
│
├── scripts/
│   ├── 01-system-health-check.sh
│   ├── 02-server-network-audit.sh
│   ├── 03-service-manager.sh
│   ├── 04-storage-manager.sh
│   └── 05-backup-manager.sh
│
└── screenshots/

3.0 Script Overview

The repository contains five modular Bash utilities that automate common DevOps operational tasks across Ubuntu Linux servers and AWS EC2 environments. Each script is designed to perform a specific operational function while remaining independent, allowing engineers to execute only the utilities required for a particular maintenance or troubleshooting task.

Script Primary Responsibilities Purpose
01-system-health-check.sh Performs a comprehensive health assessment by collecting system information including hostname, operating system, kernel version, CPU utilisation, memory and swap usage, storage utilisation, system load, uptime, logged-in users, running processes, and directory structure. Provides a quick operational overview of the server's health before troubleshooting, maintenance, or deployments.
02-server-network-audit.sh Audits network configuration by collecting EC2 metadata, public and private IP addresses, network interfaces, routing tables, DNS configuration, internet connectivity, listening ports, and time synchronisation status. Validates network connectivity and infrastructure configuration while assisting with troubleshooting and deployment verification.
03-service-manager.sh Monitors and manages critical DevOps services by checking service status, runtime engines, listening ports, Docker containers, application health, and optional restart operations. Simplifies routine service administration and enables rapid identification and recovery of failed services.
04-storage-manager.sh Analyses filesystem utilisation, inode usage, storage devices, large directories and files, Docker storage, symbolic links, backup archives, core dumps, and optional cleanup operations. Helps prevent storage-related incidents by identifying capacity issues and safely reclaiming disk space.
05-backup-manager.sh Creates, verifies, restores, rotates, and manages compressed backup archives while maintaining backup integrity and retention. Provides a reliable backup and recovery solution for protecting infrastructure configurations and application data.

4.0 Recommended Execution Order

The scripts are organised according to real-world operational workflow followed by DevOps Engineers during routine infrastructure maintenance and troubleshooting.

  1. 01-system-health-check.sh – Verify the overall health and operational status of the server.
  2. 02-server-network-audit.sh – Validate network configuration, connectivity, routing, DNS, and listening ports.
  3. 03-service-manager.sh – Inspect and manage critical services and applications running on the server.
  4. 04-storage-manager.sh – Analyse storage utilisation, filesystem health, and perform optional cleanup tasks.
  5. 05-backup-manager.sh – Create, verify, restore, and manage backups after confirming the server is healthy.

5.0 Prerequisites

Before using these scripts, ensure the following requirements are met:

  • Ubuntu Linux 22.04 LTS or later.
  • Bash shell.
  • Sudo privileges for operations that require elevated permissions.
  • Internet connectivity for external network validation and package installation (where applicable).
  • AWS EC2 instance (recommended for scripts that utilise EC2 metadata services).
  • Common Linux utilities such as curl, systemctl, ip, ss, tar, find, and df.

6.0 Quick Start

Clone the repository:

https://github.com/heyohjayy/server-operations-toolkit.git
cd devops-automation-toolkit

Make all scripts executable:

chmod +x scripts/*.sh

Run any script:

./scripts/01-system-health-check.sh

7.0 Script Documentation

The following sections describe each utility included in the toolkit, its purpose, usage, and example output. Validation screenshots are provided to demonstrate successful execution of each script.

7.1 System Health Check

The 01-system-health-check.sh utility performs a comprehensive assessment of the current server state by collecting critical operating system, hardware, resource utilisation, and process information. It is designed to provide DevOps Engineers with a quick operational overview before performing deployments, troubleshooting, maintenance, or infrastructure changes.

The report includes:

  • Hostname and system information
  • Operating system and kernel version
  • CPU architecture and utilisation
  • Memory and swap usage
  • Disk utilisation
  • System load averages
  • Logged-in users
  • Top CPU and memory-consuming processes
  • Directory tree overview

Usage

./scripts/01-system-health-check.sh

Validation

System Health Report

System Health Check - Part 1

System Health Summary

System Health Check - Part 2

7.2 Server Network Audit

The 02-server-network-audit.sh utility performs a comprehensive network audit by collecting infrastructure inventory and validating the network configuration of Ubuntu Linux servers and AWS EC2 instances. It is designed to help DevOps Engineers quickly assess network connectivity, verify cloud metadata, inspect routing configuration, and identify potential networking issues during troubleshooting or deployment activities.

The audit includes:

  • Hostname and operating system information
  • CPU and memory information
  • AWS EC2 instance metadata (IMDSv2)
  • Public and private IP address discovery
  • Network interface inspection
  • DNS configuration
  • DNS resolution tests
  • Internet connectivity validation
  • Routing table inspection
  • Listening ports audit
  • System time synchronisation
  • Consolidated network audit summary

Usage

./scripts/02-server-network-audit.sh

Validation

Server and Network Inventory

Server Network Audit - Part 1

Network Configuration and Connectivity Validation

Server Network Audit - Part 2

Network Audit Summary

Server Network Audit - Part 3

7.3 Service Manager

The 03-service-manager.sh utility audits, monitors, and manages critical DevOps infrastructure services running on Ubuntu Linux servers. It enables DevOps Engineers to verify service availability, inspect runtime engines, validate listening ports, restart individual services, and automatically recover failed services through a single operational utility.

The utility supports the following operations:

  • Audit the health status of all configured DevOps services
  • Check the status of a specific service
  • Validate runtime engines and listening ports
  • Restart individual services
  • Automatically restart failed services
  • Generate a consolidated service summary and status matrix

Usage

Display the status of all configured services:

./scripts/03-service-manager.sh --status

Display the status of a specific service:

./scripts/03-service-manager.sh --status <service_name>

Restart a specific service:

./scripts/03-service-manager.sh --restart <service_name>

Automatically restart all failed services:

./scripts/03-service-manager.sh --restart-failed

Validation

DevOps Service Status Report

Service Manager - Status Report

Service Summary and Pipeline Status Matrix

Service Manager - Summary

7.4 Storage Manager

The 04-storage-manager.sh utility performs comprehensive storage auditing and safe disk cleanup operations on Ubuntu Linux servers. It enables DevOps Engineers to monitor filesystem health, identify storage bottlenecks, locate unnecessary disk consumption, and reclaim storage through controlled cleanup operations suitable for both manual administration and automation workflows.

The utility supports the following operations:

  • Audit filesystem health and disk utilisation
  • Monitor inode usage
  • Display AWS block device mappings
  • Identify the largest directories and files
  • Detect oversized log files
  • Identify deleted-but-open (ghost) files
  • Report Docker storage usage
  • Detect broken symbolic links
  • Discover backup archives
  • Detect core dump files
  • Verify filesystem write capability
  • Perform safe storage cleanup with optional non-interactive execution
  • Generate a consolidated storage management summary

Usage

Run a comprehensive storage audit:

./scripts/04-storage-manager.sh --check

Run a safe storage cleanup:

./scripts/04-storage-manager.sh --cleanup

Run a non-interactive storage cleanup:

./scripts/04-storage-manager.sh --cleanup --force

or

./scripts/04-storage-manager.sh --cleanup -f

Validation

Storage Management Report

Storage Management Report

Filesystem Analysis and Storage Audit

Filesystem Analysis

Storage Management Summary

Storage Management Summary

7.5 Backup Manager

The 05-backup-manager.sh utility provides a complete backup management solution for Ubuntu Linux servers. It enables DevOps Engineers to create timestamped backup archives, verify backup integrity, restore archived data, manage backup retention, and maintain a reliable backup lifecycle suitable for production environments and automation workflows.

The utility supports the following operations:

  • Create compressed, timestamped backup archives
  • Verify backup archive integrity
  • Generate SHA-256 checksum files
  • Restore backup archives
  • List available backup archives
  • Rotate expired backups based on a configurable retention period
  • Maintain backup operation logs
  • Support custom backup destinations
  • Support non-interactive execution for automation and CI/CD pipelines

Usage

Create a backup:

./scripts/05-backup-manager.sh --create <source_path>

Create and verify a backup:

./scripts/05-backup-manager.sh --create <source_path> --verify

Create a backup in a custom destination:

./scripts/05-backup-manager.sh --create <source_path> --destination <backup_directory>

Restore a backup archive:

./scripts/05-backup-manager.sh --restore <backup_file.tar.gz> <target_directory>

List all available backups:

./scripts/05-backup-manager.sh --list

Rotate expired backups:

./scripts/05-backup-manager.sh --rotate <days>

Run backup rotation without confirmation:

./scripts/05-backup-manager.sh --rotate <days> --force

Validation

Backup Creation and Verification

Backup Creation

Backup Summary

Backup Summary

Backup Restore and Rotation

Backup Restore and Rotation

Available Backup Archives

Available Backups

8.0 Summary

The Server Operations Toolkit demonstrates how common infrastructure management tasks can be standardised and automated using modular Bash scripts. Each utility focuses on a specific operational responsibility while following a consistent command-line interface, making the toolkit easy to understand, extend, and integrate into existing workflows.

Whether performing routine health checks, auditing network configurations, managing critical services, analysing storage utilisation, or protecting infrastructure through automated backups, the toolkit provides practical examples of production-oriented Linux automation and operational best practices.

These utilities are designed using real-world DevOps workflows and can be adapted for development, testing, staging, and production environments with minimal modification. The scripts may also serve as reusable building blocks for CI/CD pipelines, scheduled maintenance tasks, configuration management, and day-to-day infrastructure operations.

Contributions, enhancements, and suggestions are welcome as the toolkit continues to evolve with additional automation utilities and operational capabilities.

About

A production-ready collection of Bash automation utilities for DevOps engineers to audit infrastructure, monitor services, analyse storage, validate networking, and automate backup operations on Ubuntu Linux servers.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages