Philipp Guldimann

Philipp Guldimann

AI Harness Engineer · Zürich, Switzerland

Open to AI engineering roles — Zürich or remote

I build the scaffolding that makes AI systems survive production — evaluation, agent reliability, and the data pipelines underneath.

At a Zürich AI startup I built LLM evaluation infrastructure — 23 evaluators across 10+ models — that cut QA cycles from two days to four hours. Most recently I built multi-jurisdiction legal AI pipelines processing roughly one million documents across three countries. My roots are in trustworthy AI research: I am joint first author of COMPL-AI, the first technical interpretation of the EU AI Act and an open benchmarking suite built on it. I hold an MSc in Machine Intelligence from ETH Zürich, and I care about pragmatic, scalable systems that ship and prove their value with metrics.

  • LLM evaluation
  • Agent reliability
  • Data pipelines
  • Evals & MLOps
  • Platform engineering
  • Trustworthy AI
  • Python / TypeScript
  • AWS / Azure

Writing

I write about AI harness engineering — agent loops, tool schemas, context management, evals — and where they break in practice.

Experience

Data Engineer — Omnilex

  • Built ingestion and transformation pipelines for legal content across 3 jurisdictions (Switzerland, Germany, Austria), processing ~1M documents from APIs, scraping outputs, and bulk sources.
  • Developed TypeScript-based data workflows for normalization, citation-aware chunking, embeddings, classification, and entity extraction.
  • Contributed to RAG-ready indexing and Azure-based search infrastructure for precise, traceable legal AI responses.
  • Introduced data contracts and validation checks that caught 50K+ duplicate entries before they reached production.

Machine Learning Engineer — LatticeFlow AI

  • Built v0 evaluation infrastructure for a new AI product from scratch: 23 evaluators, integrations with multiple data and chat-model providers, assessing 10+ LLMs.
  • Automated evaluation workflows for QA and regression testing, reducing review cycles from ~2 days of manual inspection to ~4 hours.
  • Extended evaluators with targeted dataset generation to increase coverage across model behaviours and failure modes.

Research

COMPL-AI — Benchmarking LLM Compliance with the EU AI Act

Joint first author (equal contribution) and lead author on COMPL-AI — the first technical interpretation of the EU AI Act, mapping its six ethical principles onto 27 concrete benchmarks and evaluating 12 prominent LLMs against them. Worked on benchmark design, evaluation pipelines, and model integration via Hugging Face Transformers. The framework is positioned as a reference point for the EU’s GPAI Code of Practice.

Read the paper (arXiv:2410.07959) · Code on GitHub

Education

MSc Computer Science — Machine Intelligence, ETH Zürich

Thesis (top grade): Speech Recognition for Children with Congenital Disorders Using Adaptive Methods — adapting Whisper to non-normative child speech from a single speaker.

Read the case study →

BSc Computer Science, ETH Zürich

Thesis (top grade): Detecting Disinformation on Twitter Targeting Non-Profit Organisations, in collaboration with the ICRC. Thesis (PDF)

Technical skills

   
Languages Python, TypeScript / JavaScript, SQL
LLM / AI LLM evaluation, RAG, embeddings, Hugging Face Transformers, MCP, PyTorch, LoRA / PEFT
Data & orchestration Dagster, PostgreSQL, pgvector, OpenSearch, Azure AI Search, Delta Lake, NATS
Backend & cloud FastAPI, NestJS, Next.js, Node.js, Docker, Kubernetes, Terraform, Argo CD, AWS, Azure, CI/CD

Earlier projects

Computational Intelligence Lab — Text Classification

Report (PDF)

Contact

I'm currently open to AI engineering roles — evaluation infrastructure, agent reliability, or the data platforms underneath — in Zürich or remote. The fastest way to reach me is email.

phil.guldimann@gmail.com LinkedIn GitHub