Technical Enablement & AI Learning | Troubleshooting, Search/RAG, Developer Education
I build tools and learning experiences that help engineers understand complex technical problems, narrow them down systematically, and use AI without giving up judgement.
My background combines 13+ years in technical support and application engineering — including five years at Elastic — with hands-on work in AI-assisted development, search/retrieval, RAG, evaluation, observability, and knowledge workflows.
The common thread across my work is turning knowledge into capability: helping people move from vague symptoms and copied answers toward evidence, diagnosis, practice, feedback, and independent problem-solving.
I am particularly interested in helping engineers learn to:
- turn vague symptoms into reproducible technical problems
- separate evidence from assumptions
- form and test useful hypotheses
- navigate unfamiliar systems and codebases
- use documentation and knowledge bases as evidence, not recipes
- recognise when a problem should be escalated
- use AI coding tools while retaining technical understanding
- verify AI-generated answers and implementations rather than trusting plausibility
I prefer AI systems that make evidence visible, measure quality, and keep humans in control where judgement matters.
A hands-on incident lab for practising diagnosis, observability, safe fixes, regression testing, and durable technical communication.
The project turns recurring support situations into reproducible scenarios and follows a practical learning loop:
observe → hypothesise → test → fix → verify → explain
What it demonstrates
- systematic debugging and incident reproduction
- CI checks, regression tests, browser E2E tests, and performance smoke tests
- OpenTelemetry-based observability and trace debugging
- Docker-based local tooling
- customer-facing runbooks and evidence-driven troubleshooting
- AI coding workflows encoded as reusable Claude Code skills
Production-shaped RAG system over technical documentation using Elasticsearch hybrid retrieval, reranking, citations, feedback, evaluation, and observability.
The central question is not just can the model answer? but does the system have enough evidence to justify an answer?
What it demonstrates
- hybrid BM25 + vector retrieval and reranking
- explicit insufficient-evidence behaviour
- citation validation and human review
- evaluation separated across retrieval, generation, and system behaviour
- MCP tools over the same retrieval/generation core
- FastAPI, PostgreSQL, Docker Compose, CI, unit/integration tests, and OpenTelemetry
Operational knowledge-quality system for finding likely duplicate or overlapping technical articles, clustering evidence, and supporting editorial review.
The design principle is deliberate: similarity is evidence, not a merge decision. Two articles can look alike while representing different versions, environments, failure modes, or user intents. Automation surfaces candidates and evidence; reviewers retain ownership of the editorial decision.
What it demonstrates
- knowledge quality and semantic similarity
- human-in-the-loop AI
- evaluation against recorded human decisions
- safe agent workflows and explicit feature gating
- provenance, auditability, and knowledge governance
- FastAPI + React/TypeScript + Elasticsearch
Interactive technical quiz focused on engineering judgement across Elasticsearch, distributed systems, observability, and resilience.
Each problem includes explanations and evidence from official documentation. The aim is to turn technical knowledge into active practice rather than passive documentation.
→ elasticsearch-resilience-quiz
Cross-platform learning application for German practice across TestDaF, workplace communication, and negotiation.
I use it to explore:
- conversational and voice-based learning
- personalised practice and learner progress
- progressive feedback instead of answer dumping
- local-first/private learning workflows
- how AI can support practice without replacing the learner's own thinking
Testing: pytest, integration tests, regression tests, Playwright/browser E2E, smoke/performance tests
Delivery: GitHub Actions, lint/typecheck/build/test gates
Runtime: Docker and Docker Compose
APIs: FastAPI, HTTP integrations, structured errors, health checks
Observability: OpenTelemetry, structured logging, tracing, metrics
Reliability: explicit failure paths, fallbacks, reproducible scenarios, rollback-aware thinking
AI engineering: RAG evaluation, grounding, human review, MCP, agent workflows
Search: Elasticsearch, BM25, vector search, hybrid retrieval, reranking, relevance evaluation
- elastic-product-search-lab — measurable search relevance with Precision@5, MRR@10, nDCG@10, and latency gates.
- elastic-search-policy-control-plane — deterministic search policies and explainable query execution.
- elastic-repo-inventory — provenance-aware technical retrieval and version-aware search.
- elastic-ai-search-decision-lab — documentation findability evaluated with practitioner questions and relevance metrics.
I am particularly interested in moving from:
documentation → retrieval → answers
toward:
knowledge → diagnosis → practice → feedback → evaluation → mastery
My current focus is not building another chatbot. It is exploring how AI can help technical people practise diagnosis, receive useful feedback, and develop transferable problem-solving skills.
Areas I want to keep exploring:
- AI troubleshooting simulations
- adaptive technical tutoring with progressive hints
- knowledge-to-learning pipelines
- learner-state and misconception modelling
- evaluation of AI tutoring quality
- conversational simulation
- small immersive/WebXR learning experiments
I prefer systems that:
- make complex knowledge understandable
- expose evidence rather than hide it
- measure quality instead of assuming it
- keep humans in control where judgement matters
- turn recurring problems into reusable knowledge and tools
- help people become more independent rather than more dependent on the tool


