We’re looking for a creative, experienced data engineer based in London to join our fintech startup at the ground floor. If you’re excited about independence, impact, and contributing to the battle against scams and financial crime, read on to learn more ⬇️
About the role đźŽ
Banks approve or reject payments with almost no context. We're building the intelligence layer that changes that — running real-time investigations on payments before they clear. You'll be our first data engineering hire, responsible for turning the rich but fragmented data we generate — payment information, payment context, and linked fraud outcomes — into a well-structured, governed, and queryable data platform. This data is becoming a product asset, not just an operational by-product, and you'll shape how it compounds into a strategic advantage for our customers.
What you’ll work on
- Design and build the data platform around our proprietary fraud and payment dataset — defining how data is modelled, where it lives, and how it flows between systems
- Consolidate and rationalise a mixed data estate spanning BigQuery, AlloyDB, Firestore, Elastic, PostHog, and a range of real-time API and scraped sources into a coherent and scalable architecture
- Build and maintain reliable, observable data pipelines that transform raw investigation data (both structured and unstructured) into forms suitable for analytics, feature engineering, and commercial data products, where outputs can be traced back to source evidence.
- Mature our data governance, security, and usage controls around highly sensitive financial data — lineage, access control, PII handling, and auditability
- Explore and build graph-oriented or relationship-based views of the data to surface the network patterns inherent in fraud
- Collaborate with ML engineers and product to close feedback loops: ensuring fraud labels, outcomes, and enrichment data flow back into the systems that need them
- Ship reliable code with strong emphasis on security and observability
- AI: yes, experimentation, endlessly. Experiment with and deploy specialised AI agents — and build the retrieval, grounding, and data plumbing needed to make LLM-powered investigation flows reliable in production, not just impressive in demos.
- Design retrieval pipelines for LLM-powered workflows: ingestion, chunking, metadata enrichment, indexing, retrieval, provenance, and evaluation.
- Team: Understand our customers to drive the right planning and prioritisation, and engage in effective code reviews (on both sides of the table!)
This is not just a “keep the pipelines running” role: you’ll be shaping and building the foundations of how we commercialise our data assets to create value for the business and our customers.
Why this is hard
The data that makes fraud detection work is messy by nature. Labels are noisy and delayed — a payment flagged today might not be confirmed as fraud for weeks. The evidence is scattered across structured databases, unstructured documents, third-party API responses, and scraped sources, with no single schema to rule them all. You need to make this data reliably queryable for both real-time scoring and retrospective analytics, while enforcing strict governance on data that is both commercially sensitive and regulated. And you're building from scratch: there's no existing data team, no mature warehouse to inherit — just a rich, valuable, and currently underused data asset waiting for the right engineer to give it structure.
Our stack today:
- Data stores: BigQuery, AlloyDB, Firestore, Elastic