Imdad Ali Khan

Senior Ruby on Rails & AI Product Engineer - I build and run AI products end to end.

I take products from empty repo to production and keep them running: LLM pipelines, vector search, backend, infrastructure, and deployment, with async-first communication and structured check-ins.

Let's Talk View LinkedIn

About

I am Imdad Ali Khan, a senior product engineer with 18+ years building and operating production systems across web, backend, mobile, and cloud infrastructure. My core is Ruby on Rails, PostgreSQL, and owning systems in production.

Over the last two years I have designed, built, and run two AI products as the only engineer, from empty repo to production:

OwlBrief is a regulatory and enforcement intelligence platform. It watches 150+ official sources, uses a multi-stage LLM pipeline to decide which changes actually bind anyone, and publishes structured decision briefs for compliance, trade, and risk teams every day.

thalo, which I co-founded, is a global recruitment marketplace and CRM for manpower agencies. It has natural-language candidate search on pgvector, an LLM CV-parsing pipeline, and a WhatsApp AI assistant that handles CV intake, job posting, and candidate search inside the chat. It went from empty repo to production in about two months.

I make decisions from measurements, write tests as a matter of course, and use AI coding tools with the judgement to keep them inside the architecture. I work best with teams that value clear priorities, async-first communication, structured check-ins, and ownership over micromanagement.

Projects

thalo — global recruitment marketplace

A global recruitment marketplace and CRM for manpower agencies, with AI search and a WhatsApp assistant.

Co-founder & sole engineer · Jul 2026 – present

Tech: Rails 8 · Hotwire · PostgreSQL 16 + pgvector · Solid Queue · OpenAI · WhatsApp Cloud API · Kamal 2 · Cloudflare

What it is

A two-sided marketplace where anyone can post a CV or a job, for any role, in any country. Employers search candidates with one natural-language box and pay only to unlock contact details. On top sits a recruitment CRM for agencies: a private candidate book, bulk import, pipelines, document tracking, and client shortlists. Built from an empty repo to production in about two months.

Key contributions & impact

  • Semantic search: Natural-language candidate and job search on OpenAI embeddings in pgvector with HNSW indexes, a similarity floor, cached query vectors, and trigram-indexed keyword paths.
  • LLM CV parsing: Reads PDF, DOCX, and image CVs into a structured profile with structured outputs, a retry ladder, stall detection, a parse cache, and per-call cost tracking. It also turns photographed job posters into draft postings.
  • WhatsApp AI assistant: Built directly on Meta's Cloud API: a signature-verified webhook with an idempotency ledger, passwordless sign-in, reverse-OTP verification, and a tool-calling LLM assistant for CV intake, job posting, and candidate search.
  • Multi-tenant CRM: Isolated each agency's candidate book with one tenancy column and scopes instead of a re-architecture. Bulk CV/ZIP/spreadsheet import with dedupe, pipelines with an append-only audit trail, and document-expiry tracking.
  • Metered billing: Per-company contact credits with a consent-based access model, row-locked spending that can't overspend under concurrency, UPI reconciliation, and tiered plans with trials.
  • Production: Two Hetzner servers on a private network behind Cloudflare, deployed with Kamal 2 and GitHub Actions, monitored by Sentry with no personal data sent, and backed by ~6,500 automated tests.

OwlBrief — know what changed, and what to do.

Regulatory and enforcement intelligence for compliance, trade, and risk teams.

Founder & sole engineer · Jun 2024 – present

Tech: Rails 8 · Hotwire · OpenAI · PostgreSQL · Redis · Sidekiq · Docker Swarm · Cloudflare

What it is

A production AI system that watches 150+ official sources (sanctions authorities, central banks, securities and trade regulators) and decides which items are executed, binding changes. Each one becomes a structured decision brief: what changed, why it matters, who must act, and which deadlines to watch.

Key contributions & impact

  • Multi-stage LLM pipeline: A regex pre-filter, four-gate LLM triage, and JSON-schema-validated summarisation with revision loops, versioned prompt caching, and tiered model selection to match cost to each task.
  • Ingestion at breadth: 40 custom regulator scrapers (OFAC, FinCEN, MAS, HKMA, RBI, SAMA, CBUAE, UN Security Council and more) alongside tiered RSS ingestion, PDF extraction, and deduplication.
  • Search and personalisation: Embedding-based semantic search, a "Read next" recommender that suppresses duplicate coverage, and user-vector personalisation.
  • Performance: Cut semantic search latency 18× (1,540 ms → 87 ms) through profiling-led caching of the vector matrix.
  • Clarify with AI: Grounded, streaming Q&A on every brief, with PII masking and tiered monthly quotas tied to billing and trials.
  • Full ownership: Backend, infrastructure, CI/CD, email delivery, SEO and AI-answer-engine visibility, and analytics that separate human visits from crawler traffic.

Research Hub API — High-performance backend platform

Production API for large-scale research aggregation

Tech: FastAPI · Python · MySQL · Vector DB · Redis · Docker · Terraform · Ansible · CI/CD

What it is

A production backend powering large-scale research ingestion, semantic search, and retrieval workflows with a strong focus on performance and reliability.

Key contributions & impact

  • High-performance services: Designed backend services handling large datasets with consistently low response times.
  • Semantic retrieval: Built vector database pipelines to support semantic search and retrieval use cases at scale.
  • Infra + delivery: Owned infrastructure provisioning, deployments, and delivery using Terraform, Ansible, and automated CI/CD pipelines.

Research Hub Mobile — Offline-first mobile application

Cross-platform research workflows for low-connectivity environments

Tech: React Native · Expo · Offline Sync · CI/CD

What it is

A cross-platform mobile application supporting real-time collaboration and offline-first research workflows.

Key contributions & impact

  • Single codebase: Built a shared iOS/Android codebase with clean, maintainable API integration.
  • Offline-first reliability: Implemented robust sync patterns to ensure usability under poor network conditions.
  • Performance tuning: Optimized performance and data handling for field usage and low-connectivity scenarios.

Unicircles — Scalable community platform

Professional networking platform built for reliability

Tech: Ruby on Rails · MySQL · Redis · AWS S3 · Docker · CI/CD

What it is

A professional networking platform designed for performance, scalability, and long-term operational stability.

Key contributions & impact

  • Reliability: Maintained 99.99% uptime using zero-downtime deployment practices.
  • Scaling: Scaled to thousands of active users through caching and background processing.
  • Product capabilities: Built secure content delivery and real-time notification systems.

What I Deliver

AI-driven backend systems

Ruby on Rails 8, Hotwire, PostgreSQL, FastAPI, Python

Design and build production backend systems that power AI-driven products, marketplaces, and data pipelines.

Production LLM features

OpenAI API, structured outputs, tool calling, JSON-schema validation

Build multi-stage LLM pipelines, document parsing, and chat assistants with versioned prompts, model tiering, and cost caps.

Search and recommendations

Embeddings, pgvector (HNSW), pg_trgm, vector databases

Ship semantic search, recommenders, and personalisation, tuned by profiling rather than guesswork.

Reliable infrastructure and deployments

Docker, Kamal 2, Docker Swarm, CI/CD, Cloudflare, Hetzner, AWS, Terraform, Ansible

Own infrastructure end to end with automated deployments, zero-downtime releases, and monitoring.

Performance and background processing

Redis, Sidekiq, Solid Queue, HTTP and edge caching

Improve latency, reliability, and throughput using caching strategies and asynchronous processing.

Integrations and payments

WhatsApp Cloud API, Web Push, transactional email, OAuth, usage metering

Connect products to the channels users live in, and build billing that stays correct under concurrency.

End-to-end product ownership

Architecture, delivery, operation, go-to-market

Take products from initial design to deployment and long-term operation without handoffs.

Let's Work Together

Available for full-time engagements with end-to-end ownership and clear outcomes.

Email me the problem, expected timeline, and engagement type. I will reply with a plan and next steps.