Skip to content
View Lucas-Arteaga's full-sized avatar

Block or report Lucas-Arteaga

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Lucas-Arteaga/README.md

Lucas Arteaga

Petroleum Engineering @ UBA · Technical Contract Analyst & Expeditor @ Valbol Worcester

I understand the economics behind every technical decision — and I optimize operating flows with data.

Portfolio site LinkedIn Email Buenos Aires, Argentina Open to relocate to Neuquén


🛠️ What I do

Fifth-year Petroleum Engineering student at Universidad de Buenos Aires, working full time at an API valve manufacturer that supplies upstream operators. I review engineering specifications for a living, I know what a specification deviation costs once it reaches the shop floor, and I automate the analysis around it.

Where I sit: between hard petroleum engineering — API 6A/6D, ASME, ASTM, well and production data — and operating data: Python, SQL, machine learning and workflow automation.

🔧 Technical compliance Review and validate engineering documentation for ball, control, butterfly and retention valves against API 6D, API 6A, ASME B16.34, ASME B16.5 and ASTM — detecting data sheet deviations before manufacturing starts, not after.
🚚 Expediting & delivery risk Coordinate critical deliveries with the shop floor, anticipate blockers, and report status to operators with numbers instead of optimism.
📊 Well & production data Certified in OpenWells, Data Analyzer and PROFILE (Halliburton Landmark). Production analytics, GOR/RGP modelling, well KPIs, reservoir characterization.
⚙️ Automation & AI Python and SQL pipelines, scikit-learn models, n8n workflows, API integrations, Docker. Certified in AI Automation, advanced level.

📈 Proof, not adjectives

270 certified hours across petroleum data science, industrial operations, industrial software and technical communication
174,815 raw well-month records processed in the project awarded a special mention by Fundación Sadosky & Fundación YPF
125,018 records surviving a documented cleaning policy (non-physical production removed, p99.5 outlier cap)
5 regression models compared under 5-fold cross-validation, with a cost/benefit call, not just a leaderboard

🧭 Tech stack

Petroleum & operations

API 6A / 6D ASME B16.34 / B16.5 ASTM OpenWells Data Analyzer PROFILE Aspen HYSYS QROD E&P concepts Reservoir & production

Data & code

Python pandas NumPy scikit-learn SQL Matplotlib Seaborn Jupyter

Automation & delivery

n8n FastAPI Docker REST APIs Git SAP ERP Excel + VBA


📂 Selected work

🏅 GOR / RGP regression comparison — special mention, Fundación Sadosky & Fundación YPF

The 2025 course project (48 h) that earned a special mention. I processed 174,815 well-month records and kept 125,018 after applying a documented cleaning policy — 48,533 rows were non-physical (zero or negative production). The target was heavily skewed, so it was log-transformed before modelling. Five models were compared under 5-fold cross-validation: a linear baseline, decision trees at two depths, and random forests at two sizes.

The decision that mattered was not the leaderboard. The 100-tree forest reached R² 0.57; the 50-tree forest reached 0.56 and trained in half the time — 266 s against 519 s. In an operating environment, that is what decides which model actually gets deployed.

Scale 174,815 raw records → 125,018 after cleaning
Target GOR / RGP, log-transformed
Models linear regression · decision tree (depth 5 and 10) · random forest (50 and 100 trees)
Validation 5-fold cross-validation · MAE, RMSE, R²
Result best R² 0.57 · deployed choice: 0.56 at half the training cost
Limits disclosed one month per year (no decline curves) · no bottom-hole pressure · no petrophysics

Why there is no repository link here: the dataset included operator production and drilling records, and it is not mine to publish. The public, reproducible version — rebuilt on Argentina's open well-production dataset (CC-BY-4.0) — is the project in progress below. Not publishing another company's operational data is part of the job.

CRISP-DM data mining over official Argentine per-well production data: business filters, percentile-based categorisation, a 20×20 Self-Organizing Map with KMeans over the SOM weights, and cluster profiling with radar charts.

Python · pandas · MiniSom · scikit-learn · Jupyter

🔨 In progress — well-production-analytics

A reproducible production analytics pipeline over Argentina's official open well-production dataset (Secretaría de Energía, CC-BY-4.0): validated ingestion of multi-year files, data-quality reporting, basin and operator profiling, and GOR modelling. Public data, so anyone can run it — and nothing confidential to hide behind.


🎓 Certifications

Certification Issuer Year Hours
Gas Plant Operator — certificate 810/24 Instituto Tecnológico de la Patagonia 2024 96
Data Science for Petroleum Engineers: first steps into AI — special mention for Comparison of regression models for GOR prediction Fundación Sadosky & Fundación YPF 2025 48
AI Automation career track — 13 weeks; includes the AI Automation and advanced-level courses Coderhouse 2026 48
Microsoft Excel, advanced level — certificate 12321 UTEPSA Postgrado 2025 30
SAP and Excel integration with script — certificate 12405 UTEPSA Postgrado 2025 24
OpenWells, Data Analyzer & PROFILE — code NWS-2026-016, publicly verifiable Next Well Solutions / Halliburton Landmark 2026 12
Public speaking and Storytelling — Res. (D) 5849/24 Universidad de Buenos Aires, Facultad de Derecho 2024 12

🧠 How I work

Four things I actually do, every time:

1. Validate before it gets built. A deviation caught on a data sheet costs a revision. The same deviation caught on the shop floor costs a scrapped part, a delayed delivery and an angry operator. The earlier the check, the cheaper the fix — that is true of data and of valves.

2. Clean the data before touching the model. In the project that earned the special mention, 48,533 of 174,815 records were non-physical — zero or negative production. Modelling on top of those would have produced a confident answer to a question nobody asked.

3. Report the cost of a decision, not just the number. The 100-tree forest scored 0.57 and the 50-tree forest scored 0.56. The second one trains in half the time. In a plant, that difference decides which one gets deployed.

4. Say what does not work. Every analysis has limits. Mine had a snapshot-only dataset, no bottom-hole pressure and no petrophysics — so the model describes wells, it does not predict them. Knowing which of the two you have is most of the job.


📬 Contact

  • LinkedIn — lucas-arteaga- (best for opportunities)
  • Site — lucas-arteaga.github.io — the same information, laid out properly
  • Email — larteaga@fi.uba.ar
  • Based in Buenos Aires · open to relocation to Neuquén (Vaca Muerta)
  • Languages — Spanish (native) · English (professional working proficiency)
Open to Petroleum Engineering, production data, well operations and technical automation roles.

🇪🇸 Leer esta página en castellano

🛠️ Qué hago

Estudiante de quinto año de Ingeniería en Petróleo (UBA), trabajando a tiempo completo en un fabricante de válvulas API que provee a operadoras de Upstream. Reviso especificaciones de ingeniería todos los días, sé cuánto cuesta una desviación de especificación cuando ya llegó al taller, y automaticé el análisis alrededor de eso.

Dónde me paro: entre la ingeniería dura — API 6A/6D, ASME, ASTM, datos de pozo y de producción — y los datos operativos: Python, SQL, machine learning y automatización de flujos.

🔧 Cumplimiento técnico Reviso y valido documentación de ingeniería de válvulas esféricas, de control, mariposa y retención bajo API 6D, API 6A, ASME B16.34, ASME B16.5 y ASTM — detectando desvíos en data sheets antes de que arranque la fabricación, no después.
🚚 Expediting y riesgo de entrega Coordino entregas críticas con planta, anticipo bloqueos y reporto estado a las operadoras con números, no con optimismo.
📊 Datos de pozo y producción Certificado en OpenWells, Data Analyzer y PROFILE (Halliburton Landmark). Analítica de producción, modelado de RGP/GOR, KPIs de pozo, caracterización de reservorios.
⚙️ Automatización e IA Pipelines en Python y SQL, modelos con scikit-learn, flujos en n8n, integración de APIs, Docker. Certificado en AI Automation, nivel avanzado.

📈 Prueba, no adjetivos

  • 270 horas certificadas entre ciencia de datos aplicada a petróleo, operación industrial, software industrial y comunicación técnica.
  • 174.815 registros pozo-mes procesados en el trabajo que recibió mención especial de Fundación Sadosky & Fundación YPF.
  • 125.018 registros que sobrevivieron a una política de limpieza documentada (se descartó producción no física y se acotaron outliers en el percentil 99,5).
  • 5 modelos de regresión comparados con validación cruzada de 5 folds, con decisión de costo-beneficio y no sólo un ranking.

🎓 Certificaciones

Certificación Emisor Año Horas
Operador de Plantas de Gas — certificado 810/24 Instituto Tecnológico de la Patagonia 2024 96
Fundamentos en Ciencias de Datos para Ingenieros en Petróleo: primeros pasos hacia la IA — mención especial por Comparación de modelos de regresión para la predicción de RGP Fundación Sadosky & Fundación YPF 2025 48
Carrera de AI Automation — 13 semanas; incluye los cursos de AI Automation y nivel avanzado Coderhouse 2026 48
Microsoft Excel, nivel avanzado — certificado 12321 UTEPSA Postgrado 2025 30
Integración de SAP y Excel con script — certificado 12405 UTEPSA Postgrado 2025 24
OpenWells, Data Analyzer y PROFILE — código NWS-2026-016, verificable públicamente Next Well Solutions / Halliburton Landmark 2026 12
Oratoria e Introducción al Storytelling — Res. (D) 5849/24 Universidad de Buenos Aires, Facultad de Derecho 2024 12

🧠 Cómo trabajo

1. Validar antes de construir. Una desviación detectada en el data sheet cuesta una revisión. La misma desviación detectada en el taller cuesta una pieza rechazada, una entrega demorada y una operadora enojada. Cuanto más temprano el control, más barata la corrección — vale para datos y vale para válvulas.

2. Limpiar los datos antes de tocar el modelo. En el trabajo que ganó la mención, 48.533 de 174.815 registros eran no físicos: producción cero o negativa. Modelar encima de eso habría dado una respuesta segura a una pregunta que nadie hizo.

3. Reportar el costo de una decisión, no sólo el número. El bosque de 100 árboles dio 0,57 y el de 50 dio 0,56, entrenando en la mitad del tiempo. En una planta esa diferencia decide cuál se pone en producción.

4. Decir lo que no funciona. Todo análisis tiene límites. El mío tenía datos de una sola foto al año, sin presión de fondo ni petrofísica — así que el modelo describe pozos, no los predice. Saber cuál de las dos cosas tenés es la mayor parte del trabajo.

📬 Contacto

  • LinkedIn — lucas-arteaga-
  • Sitio — lucas-arteaga.github.io — la misma información, mejor presentada
  • Email — larteaga@fi.uba.ar
  • Base Buenos Aires · con disponibilidad de relocalización a Neuquén (Vaca Muerta)
  • Idiomas — español (nativo) · inglés (competencia profesional)

Popular repositories Loading

  1. Lucas-Arteaga.github.io Lucas-Arteaga.github.io Public

    HTML

  2. Lucas-Arteaga Lucas-Arteaga Public

  3. well-production-data-mining well-production-data-mining Public

    CRISP-DM data mining on official Argentine oil & gas well production data — well clustering and production profiling in Python.

    Jupyter Notebook

  4. well-production-analytics well-production-analytics Public

    Production analytics over Argentina's official open per-well oil and gas dataset. Units declared once and tested, so a gas-oil ratio can never be off by 1000.

    Python