Data Engineer
We build bespoke AI solutions for legal work.
We work at the boundary of enterprise-grade software, regularly building services nobody here has built before. That makes for a dynamic environment with real opportunity to learn and requires good judgement and critical thinking.
We are looking for a Data Engineer to join our AI Acceleration Team. Day to day you will partner with AI, DevOps, and Software Engineers to build, scale, and maintain the data layer that powers every AI capability we ship. You will own the data infrastructure that powers our AI Deployment Platform, high-throughput ingestion pipelines, vector databases, audit-trail systems, and data-residency controls. This is not an application-layer AI role; you are building the infrastructure those systems depend on.
We understand that you’re unlikely to arrive with a full knowledge of legal workflows or our stack. You will have the full support of the Data Science Manager and the rest of the team as you settle into the role. Judgement, curiosity and an appetite for high levels of responsibility are all important qualities.
About the team:
This is an internal development team within Cleary that creates bespoke AI solutions for the firm utilizing our software development and data science capabilities. You will be joining the Software Development team, working on productionising and deploying the solutions.
We are a tight-knit team that values ownership, honesty, and curiosity. We work across disciplines - engineering, data science, legal, and product - to deliver meaningful impact.
Our team is 100% remote and has a strong remote-first culture that enables everyone to feel like one team and contribute meaningfully to team decisions and activities.
Main Responsibilites:
- Build and scale ingestion pipelines for unstructured data (PDFs, transcripts, logs), meeting latency SLAs and data-quality benchmarks
- Own vector database infrastructure: operate and tune multi-tenant clusters, managing indexing strategies, shard balancing, and query-latency optimisation under strict SLAs
- Instrument the LLMOps data layer: build tracing and observability infrastructure enabling the AI team to debug, optimise, and audit model behaviour at scale
- Build compliance-ready audit trails for automated compliance checks, meeting EU AI Act high-risk enforcement requirements
- Implement data-residency controls: develop BYOK pipeline models enforcing client-level data isolation, encryption, and jurisdictional routing
- Define data contracts and mentor engineers: establish versioned contracts between platform subsystems, conduct code reviews, pair-program with junior engineers, and collaborate with AI engineers and domain experts
Skills and Experience:
- Minimum of 3 years dedicated data engineering or database engineering experience, with at least 1 years scaling production systems in a cloud environment (AWS or Azure preferred)
- Excellent written and verbal communication skills
- A strong desire and focus on continued improvements and personal development
- Highly proficient in Python and SQL, with a portfolio of production-grade data pipelines, ETL jobs, or database tooling you have built and maintained
- Deep experience with at least one distributed data framework (Apache Spark, Kafka, Ray, or Airflow) used in production workloads
- Production experience scaling vector search engines (Milvus, pgvector, Qdrant, Pinecone) or massive analytical data warehouses (Snowflake, BigQuery, Redshift)
- Experience with cloud platforms (AWS or Azure preferred) for deploying and monitoring data services, including CI/CD, containerisation (Docker/Kubernetes), and observability tooling
- Demonstrated ability to design data-quality checks, schema-validation pipelines, and automated monitoring/alerting for production data systems
- Ability to own a technical domain end-to-end — scope it, build it, ship it — and mentor mid/junior engineers through code reviews and pair programming
Desirable Experience:
- Terraform, Pulumi, or other Infrastructure-as-Code tooling for managing cloud environments declaratively
- Familiarity with LLMOps tracing and observability tools (e.g. LangSmith, Arize, Weights & Biases, OpenTelemetry)
- Understanding of data governance structures, data cataloguing, lineage tracking, or metadata management frameworks
- Domain experience in legal tech, compliance tech, or other regulated industries where accuracy, auditability, and data residency are non-negotiable
- Experience designing event-driven or streaming architectures for real-time data processing
- Experience with knowledge graphs, ontologies, or semantic reasoning over structured legal data
Additional Information
- Fully remote within the UK, with occasional travel to a Cleary London office for events
- Four-day working week, with no reduction in salary; fifth day for independent study
- Salary of £70,000 to £90,000 per annum (dependent on experience)
If you meet some — not all — of the above criteria, we still encourage you to apply. We value learning ability, adaptability, and thoughtful engineering judgment above box-ticking.
If you are interested in applying, please submit a CV and short cover letter to the London Human Resources Team, LON-HR@cgsh.com.
Administrative Careers: London
- Administrative Careers: London
-
Current Opportunities
Current Opportunities
- Accounting Specialist - 12 month FTC
- Sales Director - ClearyX
- Business Development Specialist, M&A and Private Equity
- Business Development Coordinator, Capital Markets, Finance and Restructuring
- Risk Analyst
- Business Development Specialist, Capital Markets, Finance and Restructuring
- Junior Software Engineer
- Data Engineer
- Product Manager
- Privacy Notice