Portfolio of Architect L., an AI systems architect building reliability infrastructure for AI agents — memory that audits itself, retrieval that proves its answers, and evaluation harnesses. Sections: the FORGE family of systems, his thesis, credentials, and contact.

The architect's thesis

AI serves whoever directs it.
It belongs to whoever owns it.

I build the layer that makes directing it trustworthy — memory that audits its own beliefs, retrieval that proves every answer, and agents that fail closed instead of confidently wrong. Six live systems, each one proving a single property production AI actually needs.

Architect L.AI Systems Architect · Manila · GMT+8 · open to remote

The FORGE family

Not a portfolio of demos. A set of proofs — each system settles one question production AI keeps failing.

Every project below isolates a single reliability property and makes it visible, testable, and live. The point isn't that the AI answered. It's that you can see whether it can be trusted to.

Proves · provable answersLive

LEMMA

A research assistant that proves its answers. Hybrid retrieval (dense + BM25 + RRF), native citations to the exact passage, and bounded multi-hop search that decides how hard to look — with a live panel measuring its own recall, latency and cost.

RAGcitationsmulti-hopFastAPI72 tests
Proves · accountable memoryLive

ENGRAM

A memory-reliability ledger for AI agents, with a 3D observatory you fly through. What an agent believed, when it changed, what evidence triggered the revision, what decayed and was restored — WAL-hardened, proven with 24 concurrent writers losing zero records.

memoryprovenanceThree.jsSQLite
Proves · one door, every modelLive

Axiom AI

A unified gateway across Claude, GPT, Gemini and Groq — streaming, auth, usage tracking. Its Failure Contract Lab lets you click the failure path: broken streams end with a named error, not a fake success; failed calls never corrupt a session.

LLM gatewaySSEreliability30 tests
Proves · agents get testedLive

Agent Reliability Arena

An eval harness that catches what breaks agents in production: stale memory stated as truth, unsupported tool claims, answer drift, and long-horizon trend decay. Deterministic scoring — honest missing data over a fabricated leaderboard.

evalsobservabilitydeterministic
Proves · automation fails safelySource · v0.1

FlowProof

The operational half of reliability. Idempotent intake so a retried webhook creates nothing twice; bounded retries that end in an auditable dead-letter state instead of looping forever; human approval before ambiguous work runs. Every attempt in the ledger.

idempotencyaudit trailFastAPI35 tests
Proves · systems you can exploreLive

FORGE Neural Map

A real codebase rendered as a navigable 3D universe — 2,778 nodes · 7,295 connections · 157 systems — vanilla Three.js with GPU-shader layout over a live knowledge graph. Complexity you can fly through instead of read.

Three.jsWebGLdataviz
Proves · AI that compoundsPrivate · running 24/7

Project Maxima

My 24/7 personal AI partner — 219 callable tools, persistent long-horizon memory, current-truth override, and proactive pattern detection across weeks. Built and run non-stop since March; this year I took its 7,789-line core apart with a net of tests, live, without breaking it.

agentslong-horizon memorytool use
Proves · agents that self-governOpen source

LOOPKIT

A file-based operating system for self-improving autonomous loops: a charter, an append-only log, a fleet dashboard, a golden regression set, and a captions airlock so it can never post as you. Born from a documented run where its own audit caught it inflating a number and forced a correction.

agent autonomygovernanceMIT

Why I build this

The next useful AI layer isn't a bigger chat window. It's infrastructure for agents that preserve context, update beliefs when the evidence changes, and never pass stale memory off as current truth.

AI will divide people into architects — those who deliberately build with it — and fuel. I'm building to stay on the first side, and to make that side trustworthy enough to hand real work.

Most portfolios prove a candidate finished the work. Mine is built to prove I can judge whether the work is good — every system defines a failure mode, builds a repeatable pipeline, and proves progress with a credible evaluation setup.

Project Maxima stays private because it holds real, personal long-horizon memory — the thing the whole thesis is about. The public systems are how I show the engineering without exposing the data.

Read the manifesto → AI Will Not Benefit Humans

At a glance

6 live
Systems deployed
219
Tools in Maxima
3-agent
Self-governing loop
100%
Reliability-first
B.S. Information Technology — STI, 2025 CCNA ×2 — Cisco / NetAcad Sagility · US healthcare ops · trained 15+ agents · KT lead deployed to Hyderabad at 24

Open to remote AI roles

Building the layer that makes AI trustworthy. Let's talk.

AI systems, LLM agents, RAG, reliability engineering. Manila · GMT+8 · comfortable across timezones.