Job Description
Role Summary We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform: What You'll Own - Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON). - AWS-based data architecture : S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed. - Data validation frameworks : GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting. - Agent run logging & observability : designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback). - AI Factory monitoring dashboards : operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent. - ML data pipeline support : dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues. - APIs : designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools. - Data quality & testing discipline : idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production. Key Skills — Non-Negotiable (Must-Have, Strong Level) - Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting. - SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries). - AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale. - Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions." - ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built. - Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV). Key Skills — Good to Have - Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable. - FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs. - ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself. - Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats). - Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue. Skills:- Generative AI, LangGraph, ETL, databricks, Retrieval Augmented Generation (RAG) and FastAPI
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.