Data Architect + AI

Turtle Trax S.A.


Fecha: hace 17 horas
ciudad: Guadalajara, Jalisco
Tipo de contrato: Tiempo completo
Overall Stack: Databricks Lakehouse Platform, Apache Spark, Delta Lake, Unity Catalog, and modern cloud data architecture

Must Have

Lakehouse & Medallion Architecture: Expertise in designing end-to-end data architectures (Bronze, Silver, Gold layers) for reliable, production-ready pipelines

Databricks & Spark Internals: Deep understanding of distributed computing. They must know how to troubleshoot and tune large-scale Spark jobs using caching, partitioning, and broadcast joins

Delta Lake: Must understand ACID transactions, schema enforcement, time travel, and optimization operations like Z-ordering

Unity Catalog & Data Governance: Proven ability to design unified governance models for data and AI assets, including role-based access control (RBAC), row/column-level security, and data lineage

Cloud Infrastructure (AWS, Azure, or GCP): Strong grasp of the native cloud ecosystem they work in (e.g., ADLS/Entra for Azure, S3/IAM for AWS), including Virtual Network (VNet) setups and IAM roles

Coding Proficiency: Advanced SQL skills and fluency in Python or Scala

Cost Optimization & Performance Tuning: Ability to monitor DBUs (Databricks Units), right-size serverless and multi-node clusters, and implement best practices for avoiding cloud bill shock

Nice to Have

Databricks Certifications: Candidates holding valid Databricks Certified Data Architect or Databricks Certified Data Engineer Professional badges generally have a proven, up-to-date baseline of the platform's features

Generative AI & MLflow Integration: Experience building, deploying, and monitoring GenAI applications and ML models using Databricks Model Serving, Vector Search, and the Mosaic AI suite

CI/CD & DevOps Practices: Experience automating Databricks workflows using Git (Databricks Repos) and orchestration tools like dbt, Azure Data Factory, or Apache Airflow

Streaming Data: Familiarity with Databricks Structured Streaming and Auto Loader for real-time data ingestion and processing

Data Warehousing & BI: Understanding of Databricks SQL, Serverless Warehouses, and integration with downstream BI tools like Power BI

Remote

Adavenced english

Cómo postularme

Para solicitar este empleo, debe autorizarse en nuestro sitio web. Si aún no tiene una cuenta, regístrese.

Publicar un currículum