Independent engineering practice / Hamza Ahmed

I build the systems
behind the product.

Production data platforms, backend infrastructure, applied AI, and automation - designed and shipped end to end by a senior engineer who can move from architecture to implementation without losing the thread.

Selective availability for focused freelance engagements.

8+years building production systems
Millionsof records processed per day
96%pipeline infrastructure cost reduction
80%less manual support work

Engineering work delivered across environments for

GoogleNikeEpic GamesHitachi Energy

A broad engineering range, one accountable owner

Not a narrow specialist passed between five teams.

I work across the boundaries where projects usually slow down: data architecture, application code, infrastructure, third-party systems, ML and LLM workflows, observability, and delivery. The goal is not to touch every tool. It is to keep the whole system coherent.

Explore capabilities

Selected work

Systems with an architecture,
not just a screenshot.

View all eight case studies

Public reference · Data platform

GitHub ↗

MongoDB to PostgreSQL CDC with SCD2 History

A queue-first snapshot and change-data-capture service that preserves event order, durable resume state, complete SCD2 history, and a current PostgreSQL view.

MongoDBSnapshot boundarySnapshot + CDCDurable queueOrdered SCD2
PythonMongoDB change streamsPostgreSQLSCD2
View case study

Public reference · Serverless backend

GitHub ↗

Serverless Streamer Monitoring & Campaign Analytics

An event-driven AWS workflow that detects when sponsored creators go live, starts session monitoring, and records audience, chat, transcript, and mention telemetry.

Scheduled roster pollLive detectionSession lockStep Functions loopTelemetry events
AWS LambdaStep FunctionsEventBridgeS3
View case study

Public reference · Applied AI

GitHub ↗

Databricks LLM Content Pipeline with Human Review

A registry-driven workflow that screens source material, generates structured content, routes invalid output, and prepares valid results for human review.

Content registrySource aggregationDeterministic screenModel fallbackStructured generation
DatabricksPySparkDelta LakeLLM endpoints
View case study

Public reference · Generative AI

GitHub ↗

Databricks LoRA Creative Asset Pipeline

A GPU-oriented workflow for adapting SDXL to an existing visual language, generating traceable variations, filtering failures, and handing candidates to designers.

Reference assetsBLIP captionsLoRA trainingMLflow + promptsBatch generation
SDXLLoRADiffusersPEFT
View case study

The public repositories are anonymized reference implementations based on systems I designed and built. They use synthetic data and contain no client credentials, proprietary configuration, private datasets, or confidential business logic.

Selected outcomes

Engineering measured by
what changed afterward.

Results across startup, consulting, product, and enterprise environments.

01 / Infrastructure96%

lower pipeline infrastructure cost after reworking streaming compute, compression, and storage.

02 / Support automation80%

less manual response time through agentic routing, dedicated conversations, summaries, and personalization.

03 / Trust and safety87%

fewer abuse incidents after deploying a real-time moderation service into production workflows.

04 / Cloud performance45%

faster Athena queries while reducing cost by 30% through partitioning, bucketing, and compression.

Bring the problem, not a preselected stack

Build, repair, modernize, or automate.

01

Data platforms & pipelines

Warehouses, lakehouses, CDC, streaming, ETL/ELT, dbt, Airflow, Spark, data quality, history models, recovery, and migration.

Snowflake · Databricks · PostgreSQL · MongoDB · Kafka · AWS
02

Backend systems & integrations

Python, Go, Rust, and TypeScript services; APIs; serverless workflows; internal tools; SaaS connections; deployment and operational automation.

Python · Go · Rust · TypeScript · Lambda · Cloud Run
03

Applied AI & agentic workflows

LLM pipelines, RAG, model serving, text-to-SQL, support automation, human-review systems, ML workflows, and multi-agent engineering operations.

LLMs · RAG · MLflow · Databricks · Agents · MCP
04

Reliability, cost & delivery

Architecture reviews, performance tuning, failure recovery, idempotency, testing, observability, cloud cost reduction, CI/CD, and production handoff.

Terraform · Docker · Kubernetes · GitHub Actions · CI gates

How the work moves

Clear decisions.
Visible progress.
No mystery phase.

Small debugging work can stay small. Larger builds get enough structure to make architecture, assumptions, tests, and handoff explicit.

  1. 01

    Frame the real problem

    Define the business outcome, current failure mode, constraints, source of truth, and what success must prove.

  2. 02

    Design the smallest complete system

    Choose boundaries, data contracts, failure behavior, and a delivery path before adding unnecessary infrastructure.

  3. 03

    Build with evidence

    Implementation, tests, operational checks, and measurable acceptance criteria move together rather than being separate phases.

  4. 04

    Verify and hand off

    Document what exists, what was proven, how to operate it, and where the remaining risks or next decisions actually are.

About Hamza Ahmed

Engineer, builder, and founder with a physics habit of asking what the system is actually doing.

I am a Senior Software, Data & AI Engineer with eight years across startups, consulting, product teams, and enterprise client environments. I have owned systems from ingestion and orchestration through APIs, infrastructure, analytics, ML, deployment, and the workflows people use on top of the data.

Before software and data engineering, I studied physics and mathematics and worked in astrophysics research. That background still shapes the work: define the model, test assumptions, trace evidence, and make reliability part of the design.

Knownbyfew is my independent engineering practice. Martian Bee is the product company where I am combining the same disciplines into a private analytics platform.

Start with the problem

What needs to work better?

A failing pipeline, an unfinished platform, a hard integration, an AI workflow that needs production discipline, or a system that has outgrown its first architecture.

No mailing list. No sales sequence. Just a direct reply.