Petroleum Engineering @ UBA · Technical Contract Analyst & Expeditor @ Valbol Worcester
I understand the economics behind every technical decision — and I optimize operating flows with data.
📌 What I do · Proof · Work · Certifications · How I work · Español
Fifth-year Petroleum Engineering student at Universidad de Buenos Aires, working full time at an API valve manufacturer that supplies upstream operators. I review engineering specifications for a living, I know what a specification deviation costs once it reaches the shop floor, and I automate the analysis around it.
Where I sit: between hard petroleum engineering — API 6A/6D, ASME, ASTM, well and production data — and operating data: Python, SQL, machine learning and workflow automation.
| 🔧 Technical compliance | Review and validate engineering documentation for ball, control, butterfly and retention valves against API 6D, API 6A, ASME B16.34, ASME B16.5 and ASTM — detecting data sheet deviations before manufacturing starts, not after. |
| 🚚 Expediting & delivery risk | Coordinate critical deliveries with the shop floor, anticipate blockers, and report status to operators with numbers instead of optimism. |
| 📊 Well & production data | Certified in OpenWells, Data Analyzer and PROFILE (Halliburton Landmark). Production analytics, GOR/RGP modelling, well KPIs, reservoir characterization. |
| ⚙️ Automation & AI | Python and SQL pipelines, scikit-learn models, n8n workflows, API integrations, Docker. Certified in AI Automation, advanced level. |
| 270 | certified hours across petroleum data science, industrial operations, industrial software and technical communication |
| 174,815 | raw well-month records processed in the project awarded a special mention by Fundación Sadosky & Fundación YPF |
| 125,018 | records surviving a documented cleaning policy (non-physical production removed, p99.5 outlier cap) |
| 5 | regression models compared under 5-fold cross-validation, with a cost/benefit call, not just a leaderboard |
| Petroleum & operations |
|
| Data & code |
|
| Automation & delivery |
|
The 2025 course project (48 h) that earned a special mention. I processed 174,815 well-month records and kept 125,018 after applying a documented cleaning policy — 48,533 rows were non-physical (zero or negative production). The target was heavily skewed, so it was log-transformed before modelling. Five models were compared under 5-fold cross-validation: a linear baseline, decision trees at two depths, and random forests at two sizes.
The decision that mattered was not the leaderboard. The 100-tree forest reached R² 0.57; the 50-tree forest reached 0.56 and trained in half the time — 266 s against 519 s. In an operating environment, that is what decides which model actually gets deployed.
| Scale | 174,815 raw records → 125,018 after cleaning |
| Target | GOR / RGP, log-transformed |
| Models | linear regression · decision tree (depth 5 and 10) · random forest (50 and 100 trees) |
| Validation | 5-fold cross-validation · MAE, RMSE, R² |
| Result | best R² 0.57 · deployed choice: 0.56 at half the training cost |
| Limits disclosed | one month per year (no decline curves) · no bottom-hole pressure · no petrophysics |
Why there is no repository link here: the dataset included operator production and drilling records, and it is not mine to publish. The public, reproducible version — rebuilt on Argentina's open well-production dataset (CC-BY-4.0) — is the project in progress below. Not publishing another company's operational data is part of the job.
CRISP-DM data mining over official Argentine per-well production data: business filters, percentile-based categorisation, a 20×20 Self-Organizing Map with KMeans over the SOM weights, and cluster profiling with radar charts.
Python · pandas · MiniSom · scikit-learn · Jupyter
A reproducible production analytics pipeline over Argentina's official open well-production dataset (Secretaría de Energía, CC-BY-4.0): validated ingestion of multi-year files, data-quality reporting, basin and operator profiling, and GOR modelling. Public data, so anyone can run it — and nothing confidential to hide behind.
| Certification | Issuer | Year | Hours |
|---|---|---|---|
| Gas Plant Operator — certificate 810/24 | Instituto Tecnológico de la Patagonia | 2024 | 96 |
| Data Science for Petroleum Engineers: first steps into AI — special mention for Comparison of regression models for GOR prediction | Fundación Sadosky & Fundación YPF | 2025 | 48 |
| AI Automation career track — 13 weeks; includes the AI Automation and advanced-level courses | Coderhouse | 2026 | 48 |
| Microsoft Excel, advanced level — certificate 12321 | UTEPSA Postgrado | 2025 | 30 |
| SAP and Excel integration with script — certificate 12405 | UTEPSA Postgrado | 2025 | 24 |
OpenWells, Data Analyzer & PROFILE — code NWS-2026-016, publicly verifiable |
Next Well Solutions / Halliburton Landmark | 2026 | 12 |
| Public speaking and Storytelling — Res. (D) 5849/24 | Universidad de Buenos Aires, Facultad de Derecho | 2024 | 12 |
Four things I actually do, every time:
1. Validate before it gets built. A deviation caught on a data sheet costs a revision. The same deviation caught on the shop floor costs a scrapped part, a delayed delivery and an angry operator. The earlier the check, the cheaper the fix — that is true of data and of valves.
2. Clean the data before touching the model. In the project that earned the special mention, 48,533 of 174,815 records were non-physical — zero or negative production. Modelling on top of those would have produced a confident answer to a question nobody asked.
3. Report the cost of a decision, not just the number. The 100-tree forest scored 0.57 and the 50-tree forest scored 0.56. The second one trains in half the time. In a plant, that difference decides which one gets deployed.
4. Say what does not work. Every analysis has limits. Mine had a snapshot-only dataset, no bottom-hole pressure and no petrophysics — so the model describes wells, it does not predict them. Knowing which of the two you have is most of the job.
- LinkedIn — lucas-arteaga- (best for opportunities)
- Site — lucas-arteaga.github.io — the same information, laid out properly
- Email — larteaga@fi.uba.ar
- Based in Buenos Aires · open to relocation to Neuquén (Vaca Muerta)
- Languages — Spanish (native) · English (professional working proficiency)
🇪🇸 Leer esta página en castellano
Estudiante de quinto año de Ingeniería en Petróleo (UBA), trabajando a tiempo completo en un fabricante de válvulas API que provee a operadoras de Upstream. Reviso especificaciones de ingeniería todos los días, sé cuánto cuesta una desviación de especificación cuando ya llegó al taller, y automaticé el análisis alrededor de eso.
Dónde me paro: entre la ingeniería dura — API 6A/6D, ASME, ASTM, datos de pozo y de producción — y los datos operativos: Python, SQL, machine learning y automatización de flujos.
| 🔧 Cumplimiento técnico | Reviso y valido documentación de ingeniería de válvulas esféricas, de control, mariposa y retención bajo API 6D, API 6A, ASME B16.34, ASME B16.5 y ASTM — detectando desvíos en data sheets antes de que arranque la fabricación, no después. |
| 🚚 Expediting y riesgo de entrega | Coordino entregas críticas con planta, anticipo bloqueos y reporto estado a las operadoras con números, no con optimismo. |
| 📊 Datos de pozo y producción | Certificado en OpenWells, Data Analyzer y PROFILE (Halliburton Landmark). Analítica de producción, modelado de RGP/GOR, KPIs de pozo, caracterización de reservorios. |
| ⚙️ Automatización e IA | Pipelines en Python y SQL, modelos con scikit-learn, flujos en n8n, integración de APIs, Docker. Certificado en AI Automation, nivel avanzado. |
- 270 horas certificadas entre ciencia de datos aplicada a petróleo, operación industrial, software industrial y comunicación técnica.
- 174.815 registros pozo-mes procesados en el trabajo que recibió mención especial de Fundación Sadosky & Fundación YPF.
- 125.018 registros que sobrevivieron a una política de limpieza documentada (se descartó producción no física y se acotaron outliers en el percentil 99,5).
- 5 modelos de regresión comparados con validación cruzada de 5 folds, con decisión de costo-beneficio y no sólo un ranking.
| Certificación | Emisor | Año | Horas |
|---|---|---|---|
| Operador de Plantas de Gas — certificado 810/24 | Instituto Tecnológico de la Patagonia | 2024 | 96 |
| Fundamentos en Ciencias de Datos para Ingenieros en Petróleo: primeros pasos hacia la IA — mención especial por Comparación de modelos de regresión para la predicción de RGP | Fundación Sadosky & Fundación YPF | 2025 | 48 |
| Carrera de AI Automation — 13 semanas; incluye los cursos de AI Automation y nivel avanzado | Coderhouse | 2026 | 48 |
| Microsoft Excel, nivel avanzado — certificado 12321 | UTEPSA Postgrado | 2025 | 30 |
| Integración de SAP y Excel con script — certificado 12405 | UTEPSA Postgrado | 2025 | 24 |
OpenWells, Data Analyzer y PROFILE — código NWS-2026-016, verificable públicamente |
Next Well Solutions / Halliburton Landmark | 2026 | 12 |
| Oratoria e Introducción al Storytelling — Res. (D) 5849/24 | Universidad de Buenos Aires, Facultad de Derecho | 2024 | 12 |
1. Validar antes de construir. Una desviación detectada en el data sheet cuesta una revisión. La misma desviación detectada en el taller cuesta una pieza rechazada, una entrega demorada y una operadora enojada. Cuanto más temprano el control, más barata la corrección — vale para datos y vale para válvulas.
2. Limpiar los datos antes de tocar el modelo. En el trabajo que ganó la mención, 48.533 de 174.815 registros eran no físicos: producción cero o negativa. Modelar encima de eso habría dado una respuesta segura a una pregunta que nadie hizo.
3. Reportar el costo de una decisión, no sólo el número. El bosque de 100 árboles dio 0,57 y el de 50 dio 0,56, entrenando en la mitad del tiempo. En una planta esa diferencia decide cuál se pone en producción.
4. Decir lo que no funciona. Todo análisis tiene límites. El mío tenía datos de una sola foto al año, sin presión de fondo ni petrofísica — así que el modelo describe pozos, no los predice. Saber cuál de las dos cosas tenés es la mayor parte del trabajo.
- LinkedIn — lucas-arteaga-
- Sitio — lucas-arteaga.github.io — la misma información, mejor presentada
- Email — larteaga@fi.uba.ar
- Base Buenos Aires · con disponibilidad de relocalización a Neuquén (Vaca Muerta)
- Idiomas — español (nativo) · inglés (competencia profesional)