Lead Data Engineer
Firmable is the market-leading B2B sales intelligence platform in Asia-Pacific. Our competitive moat is our data: the deepest, most localised company and people dataset in every market we operate in. We've proven the model in ANZ. Now we're scaling it across Southeast Asia and the US.
Every record Firmable sells starts in a sourcing pipeline. This role architects that pipeline.
The Role
Lead Data Engineer: Sourcing designs and owns the end-to-end extraction and ETL pipeline that turns unstructured web data into the world's most accurate B2B dataset across 13 markets. You set the architectural standards: extractor patterns, proxy strategy, LLM infrastructure as production systems, agentic escalation workflows, and cost-aware orchestration.
This is not a pipeline maintenance or incremental optimisation role. You're architecting the framework other engineers ship into, building for the hardest extraction problems (anti-bot defences, JS-heavy rendering, schema drift), and owning the sourcing layer end to end.
~80% hands-on architecture and reference implementation: designing extractor frameworks, shipping the hardest extractors, building agentic pipelines, writing eval harnesses and labelled datasets. ~20% cross-functional voice: technical input on sourcing decisions with data platform, product, and analytics; setting standards for coverage, schema, and accuracy.
What You'll Own
Data Sourcing and ETL architecture
End-to-end pipeline design: extraction, normalisation, deduplication, validation, and load; plus cost, performance, and reliability of the whole layer
Sourcing layer standards: coverage, accuracy, schema design; the rule-vs-LLM standard codified so the team applies it without you
Cost and performance optimisation: cost ceilings with the token math behind them, incremental processing, recovery design, scheduling
Extraction systems and frameworks
Extractor framework: patterns, abstractions, and tooling other engineers ship into; new extractors are fast to build and reliable to run
Hard-source extractors: anti-bot defences, JS-heavy rendering, schema drift, low-quality structure; plus the proxy and IP rotation strategy behind them
Agentic extraction pipelines: rule-based triage, LLM escalation, structured-output validation, retries, human-review queues
LLM infrastructure and observability
LLMs as production systems: versioned prompts, labelled eval sets, measured precision and recall, prompt versioning you can roll back, judges debugged on real data
Eval and observability scaffolding: eval frameworks, prompt versioning, traces, drift detection when a vendor silently updates a model
Model-choice playbook: cheap models for classification, stronger models for nuanced extraction, frontier models for hard edge cases; revised as model economics shift
Skills library and orchestration
Skills library: versioned SKILL.md specs for recurring extraction patterns, invocable by engineers and agents alike
Orchestration: Airflow or equivalent patterns that scale with data volumes and source counts; dependency management, recovery, cost-aware scheduling
What We're Looking For
Must Haves
6-10 years shipping production extraction, ETL, or data pipeline systems in business-critical environments; deep web extraction at scale with experience designing around anti-bot defences, proxy architecture, JS rendering, schema drift, and recovery
Shipped LLMs inside extraction pipelines as production systems: structured outputs, versioned prompts, labelled eval sets, logged traces, judges debugged on precision/recall, drift detected on vendor model updates. You can show the repo.
Built agents and tool-calling pipelines: architected agent workflows, written SKILL.md specs others depend on, run tool-calling systems in production
Sharp judgement on rules vs. LLMs: reach for deterministic logic first, defend the call either way, and have codified the standard for a team
Expert Python: production-grade, performance-aware, comfortable with concurrency and scale; plus advanced SQL for complex transformations and performance optimisation
Extensive Airflow or equivalent: orchestration, dependency management, recovery patterns, cost optimisation at production scale
Shipped with agentic IDEs: Claude Code, Cursor, or equivalent as your default mode, with shipped extraction systems to show for it
Architecture judgement and a product mindset: when to refactor, optimise, ship, or start again; you care about coverage and accuracy landing with customers, not uptime metrics
You live and breathe AI tools. Structured outputs, evals, traces, and LLM tracing (Logfire, OpenTelemetry, similar) as your default way of working, not a productivity experiment. In 2026, this is how data engineers build extraction at scale and you need to already be doing it.
Highly Valued
Cloud data platforms: Snowflake, Redshift; AWS for pipeline deployment (Lambda, S3, ECS, Glue)
Vector databases, embeddings, or retrieval patterns for matching and deduplication
Eval frameworks like Braintrust, Promptfoo, or Inspect; LLM tracing tooling
Data quality frameworks with automated testing and anomaly detection at scale
B2B data: firmographics, people data, company registries across markets
Data privacy and compliance: GDPR, CCPA
Startup or scaleup experience where you shipped fast and owned outcomes end to end
The Environment
Firmable runs lean and ships fast: small senior teams, no layers, minimal process. Sourcing is a core competitive advantage; this role sets the pace for how fast and accurate the entire data pipeline moves. Weekly releases moving toward daily; no fixed hours, full autonomy on architecture choices.
We are an AI-native organisation. That means AI isn't a tool we reach for: it's the default operating mode. Agentic extraction pipelines, LLM-powered classification with eval sets, structured-output validation, labelled datasets versioned with prompt runs, traces logged from day one with cost and latency baked in. If you're not already working this way, this role isn't right for you.
Why This Role
Own the sourcing engine: every record Firmable sells starts in the pipeline you architect; your design shapes the accuracy and cost of every customer interaction
Greenfield AI-native architecture: the extractor framework, skills library, eval harnesses, agentic orchestration, and production LLM infrastructure are largely unbuilt; you'll architect them from first principles
Frontier technical problems: agentic extraction at production scale, LLM-as-extractor with full eval coverage, vendor model drift detection, rule-vs-LLM orchestration across 13 markets
Small team, massive leverage: your architecture reaches every Firmable customer, every day; you write the reference implementation daily and own outcomes end to end
Real career runway: this role is a bottleneck role; progress is measured by pipeline speed and accuracy, not tenure
Who This Role is NOT For
Anyone looking to optimise an existing pipeline rather than architect from first principles
Anyone who treats infrastructure as someone else's problem; sourcing architecture lives in your code every day
Anyone who sees LLMs as a parsing shortcut rather than production systems that need eval sets, versioning, and drift detection
Anyone uncomfortable shipping with minimal process or without an established playbook to follow
Anyone who treats AI as something they'll learn on the job rather than already use daily in infrastructure work
Anyone coming from a pure ops or data analytics background without shipped systems engineering at scale
- Department
- R&D - Data
- Role
- Software Engineering Lead
- Locations
- India (remote)
- Remote status
- Fully Remote