Skip to content
View PedroPauloMR's full-sized avatar

Block or report PedroPauloMR

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
PedroPauloMR/README.md

Hi 👋, I'm Pedro Paulo

Data Scientist · Data Engineer

Building robust systems to turn data into insights and help decision making.


🧠 About Me

I’m a Data Scientist and Data Engineer focused on build end‑to‑end machine learning systems that combine data engineering, statistical modeling, and cloud‑native architectures. I began my career in 2019, creating automation solutions with VBA and writing PL/SQL queries. In 2020, I transitioned fully into the Analytics field, working on data projects that blended 80% data science and 20% data engineering, which shaped my hybrid technical profile. From 2020 to 2023, I worked in the energy distribution sector, developing solutions aimed at cost reduction and operational efficiency, including:

  • optimization of operational base allocation
  • maintenance performance indicators for power distribution networks
  • meter failure prediction models
  • automated pipelines for logistics indicators
  • outage forecasting for distribution networks After 2023, I moved into consulting to work with cloud‑native applications, collaborate with diverse teams, and gain exposure to projects across multiple industries. Since 2024, I’ve been part of an international, global team, designing and deploying machine learning applications on AWS with a focus on sales‑driven solutions.

🚀 What I Work On

  • 🤖 ML Systems

    • Statistics
    • Machine-learning (classic)
    • NLP and first steps into LLMs
  • 🧱 Data Engineering

    • ETL / ELT pipelines
    • Data modeling and analytics engineering
    • Event-driven
  • ☁️ Cloud-Native & Scalable Architectures

    • Containerized services
    • Orchestrated workflows
    • Serverless applications (focus)
  • 🎓 Education & Mentorship

    • New achievements/implementations inside company (knowledge sharing)
    • Mentoring interns on my last jobs in Brazil

🛠️ Tech Stack

🧠 AI and ML Ecosystem

  • Modelling and Training · ML Pipelines
  • Deploy and MLOps with SageMaker
  • Neural Networks · Deep Learning · TensorFlow
  • Feature Engineering · scikit‑learn

🐍 Languages & Backend Frameworks


📊 Data Engineering & Analytics

  • ETL / ELT Pipelines
  • Data Modeling & Analytics Engineering
  • Batch Data Processing

🗄️ Databases and Storages


☁️ Cloud, DevOps & Infrastructure

  • Containerized Microservices
  • Cloud-Native AI Systems
  • CI/CD & Production Deployments

Some projects:

Popular repositories Loading

  1. mercado_financeiro mercado_financeiro Public

    Modelo de valuation de Fundos Imobiliários apresentado pelo time da Suno Research

    Jupyter Notebook 2 2

  2. survival_analysis survival_analysis Public

    Python project for data analysis involving survival methods of equipments, such as Kaplan-Meier and Cox Model

    Jupyter Notebook

  3. mlops-alura mlops-alura Public

    curso mlops alura

    Jupyter Notebook

  4. curso_git_cds curso_git_cds Public

    Jupyter Notebook

  5. ds_em_producao ds_em_producao Public

    Python project to understand the CRISP-DM method and apply it on Rossmann stores to predict the revenue per week (forecasting by XGBoost)

    Jupyter Notebook

  6. pa005_insiderclustering pa005_insiderclustering Public

    Python project for customer segmentation using unsupervised methods to identify the most valuables customers

    HTML