The problem
Opportunities are spread across four government sources (SAM.gov, USASpending, FPDS, SBA) with conflicting metadata and heavy overlap. Analysts spent hours per pursuit re-reading solicitations and assembling competitive intelligence by hand.
What I built
Scheduled ingestion pulls all four sources, deduplicates them in four tiers, and validates the result. Each opportunity gets a composite score that blends structured signals with vector similarity. Users ask a RAG copilot (hybrid full-text and pgvector retrieval, intent classification) and get streamed, cited answers from 50,000+ live solicitations. On request, a four-phase agent pipeline (14 coordinated agent personas, shared persistent memory) produces a full competitive-intelligence report.
How it’s built
Python / FastAPI · PostgreSQL + pgvector · Docker · Cloudflare Zero Trust, Workers, R2 · Claude and Gemini APIs · GitHub Actions, pytest, Ruff, mypy