Upwork Proposal Intelligence
Paste a job — read a grounded proposal.
Overview
Upwork Proposal Intelligence helps BD and freelancer operators write stronger Upwork proposals by finding the most relevant past company projects (~139 PCM projects) and generating a citation-grounded proposal that cites only real work and real URLs.
The system is split into two apps: a Python FastAPI backend (intelligence engine) and a Next.js frontend (operator workspace). They communicate over REST and Server-Sent Events (SSE) for live pipeline progress.
The problem: Raw project descriptions are messy prose with noisy tags (REACT vs REACT.JS). Simple keyword or single-vector search cannot pick the right past work. This is a structured-knowledge-from-unstructured-prose problem — solved by enriching projects offline and running a multi-stage retrieval + judge pipeline online.
System architecture
Technology Stack
Each technology was chosen for a specific reason. Below is what we use, and why.
Backend
| Component | Technology | Why we use it |
|---|---|---|
| Language | Python ≥ 3.11 | Strong typing, async I/O, mature LLM/ML ecosystem |
| Packaging | uv + hatchling | Fast installs; CLI entrypoints upwork and upwork-api |
| API | FastAPI + uvicorn | Async HTTP, clean OpenAPI, streaming-friendly |
| Streaming | sse-starlette | Push live stage events (decompose → retrieve → write) to the UI |
| Config | pydantic-settings | Typed .env config for all backends |
| ORM / Migrations | SQLAlchemy async + Alembic | Evolve Postgres schema safely |
| Observability | structlog, Prometheus, OpenTelemetry | Per-stage latency, tokens, correlation IDs |
Data stores
| Component | Technology | Why we use it |
|---|---|---|
| Primary DB + vectors | Postgres 16 + pgvector | One source of truth for projects, sessions, proposals, and dense embeddings |
| Full-text search | Postgres FTS / BM25 | Exact tech/tool name hits that embeddings can miss |
| Capability graph | Neo4j 5 | Projects near these capabilities — proximity beyond keyword overlap |
| Cache / cancel | Redis | Cache embeddings & job decompositions; cancel in-flight SSE runs |
| Object store | MinIO (compose) | Optional artifacts storage for local/demo stacks |
AI / ML models
| Component | Technology | Why we use it |
|---|---|---|
| Heavy LLM | gpt-oss:120b via Ollama (default) | Decompose job, judge matches, write proposal — needs strong reasoning |
| Cheap LLM | qwen2.5:7b | Lighter helper calls where full heavy model is unnecessary |
| Embeddings | Qwen3-Embedding-8B (VECTOR_DIM=1024) | Dense similarity for HyDE paragraphs and sub-questions |
| Reranker | BGE-reranker (TEI or in-process) | Cross-encode job vs candidates; precision after fusion |
| Provider swap | OpenAI-compatible clients | Change base URL / profile to switch Ollama ↔ cloud without rewriting pipeline |
Frontend
| Component | Technology | Why we use it |
|---|---|---|
| Framework | Next.js 14 (App Router) + React 18 | Fast SPA-style workspace for operators |
| Styling | Tailwind + shadcn / Base UI | Consistent, accessible controls |
| Server state | TanStack React Query | History and proposal envelope fetches |
| Live updates | @microsoft/fetch-event-source | Reliable SSE with reconnect for pipeline theatre |
| URL state | nuqs | Deep-link ?proposal= and ?review= for sharing and reload |
| Tests | Vitest | Unit coverage for stream/API helpers |
Pipeline Steps
The system has two flows: an offline flow to prepare the corpus (run once or on updates), and an online flow for every job submission.
Offline flow — prepare the corpus
Why: Persist the truth source — names, descriptions, tags, URLs from pcm_projects.json
Postgres projects table (raw rows)
Why: Fix tag noise; add domains, capabilities, stack, architecture, scale, outcomes via LLM
ProjectIntelligence JSONB in Postgres
Why: Enable dense semantic retrieval over project chunks
pgvector HNSW index on chunks.embedding
Why: Enable capability-proximity retrieval beyond keyword overlap
Neo4j capability graph from ontology YAMLs
Why: Keep knowledge current when PCM adds projects — without re-formatting or re-embedding the old corpus
New IDs only: LLM format → pgvector chunks → Neo4j Project nodes
Online flow — every job submission
Why: Aligns the job to the same schema as enriched projects; HyDE paragraphs and sub-questions improve recall
DecomposedJob: domains, capabilities, HyDE paras, sub-questions, archetype. Cached in Redis (~24h).
Why: No single channel is enough — semantic, structured, graph, and lexical each catch different misses
~top 25 candidates fused via Reciprocal Rank Fusion (k=60): dense_hyde, dense_subq, capability filter, Neo4j graph, Postgres FTS
Why: Cheap fusion is recall-heavy; cross-encoder adds precision before the expensive judge
BGE reranker scores job vs candidates → ~top 12 survivors
Why: Human-readable fit scores + rationales; quality floor for what the writer may cite
Ranked matches with domain, stack, capability, architecture, scale, trust, differentiator scores
Why: Self-check when retrieval looks weak; retry smarter before writing garbage
Broadened retrieval if confidence too low, then re-run judge
Why: Produce the actual proposal; linters block invented project IDs, URLs, and capabilities
Citation-grounded proposal text streamed via SSE proposal.delta events
Daily knowledge ingest (new projects only)
When PCM publishes a project that is not already in the vector store, the API formats it, embeds it into Postgres pgvector, and adds it to the Neo4j graph. Projects that already have rows in chunks are skipped — no re-format, no re-embed, no graph rebuild.
Download the latest project list and merge into data/raw/pcm_projects.json.
New = PCM / corpus ID with zero rows in Postgres chunks. Already-indexed IDs are never reprocessed.
Enrich only those IDs if enriched is still null.
Chunk + embed only those IDs into the chunks table (this is the vector store — there is no separate Qdrant).
Upsert Project subgraphs only for those IDs.
How to verify vector + graph
Scheduler env
| Variable | Typical | Meaning |
|---|---|---|
| PCM_SYNC_ENABLED | true | Start the background scheduler with the API |
| PCM_SYNC_INTERVAL_HOURS | 24 | Hours between ticks. Floor is 60 seconds (use 0.0167 for a 1-minute test) |
| PCM_SYNC_USE_LLM | true | Run LLM format on new IDs |
| PCM_SYNC_RUN_ON_STARTUP | false | If true, run once immediately when the API starts |
Ingest analytics
Live pgvector and Neo4j sizes, plus dated ingest runs (when new IDs were actually written). History is saved under data/ingest_reports.json whenever a run syncs or errors — skipped ticks are not logged.
Snapshot time: —
Loading live store counts…
How to Use
Step-by-step guide for operators using the workspace UI.
- 1Open the workspace
Action: Navigate to /workspace (or click Open Workspace from the landing page).
Result: You see the composer and an empty state with a sample job option.
- 2Paste the Upwork job
Action: Copy the full job description from Upwork — title, requirements, budget, skills. Paste into the composer.
Result: The job appears as a user message in the thread.
- 3Submit and watch the pipeline
Action: Click Submit. The pipeline theatre shows live progress: Decompose → Retrieve → Rerank → Judge → Write.
Result: SSE events stream stage completions and proposal text tokens in real time.
- 4Review matched projects (optional)
Action: If match review is enabled (stop_after_match), pin or drop projects before generation continues.
Result: Only selected projects are cited in the final proposal.
- 5Read the proposal
Action: Scroll through the generated proposal. Expand the audit drawer to see matched projects, scores, and confidence.
Result: A grounded proposal citing real past projects with real URLs — no invented work.
- 6Refine if needed
Action: Use the refine panel to adjust tone, length, or focus. Refinement re-runs only Stage F using cached matches.
Result: An updated proposal appended as a new turn — faster than a full re-run.
- 7Record outcome
Action: Mark the proposal as won, lost, interview, no reply, etc. from the outcome modal.
Result: Feedback stored for offline correlation — helps improve future ranking signals.
- 8Browse history
Action: Use the sidebar to filter by outcome and reopen past proposals.
Result: Full envelope restored: job, matches, proposal text, audit trail.
Query Examples
What to paste in the composer, and what the system returns for each type of job.
We're a Series A fintech startup looking for a senior full-stack engineer to build a real-time payments dashboard. Requirements: - React/Next.js front end with live-updating charts over WebSockets - Node.js + PostgreSQL backend, Stripe Connect integration - Handle 10k+ transactions/day at sub-second latency - SOC 2 awareness a plus 3-month contract with potential to extend.
- →Stage A extracts domains (fintech, payments), capabilities (real-time dashboards, Stripe integration), and tech stack signals
- →Stage B retrieves past projects with React dashboards, PostgreSQL backends, payment integrations, and WebSocket experience
- →Stage D scores matches on domain fit, stack overlap, and scale handling
- →Stage F writes a proposal citing 5–10 relevant past projects with specific URLs and outcomes
Looking for an experienced Python developer to build a production-ready RAG chatbot using FastAPI, PostgreSQL, pgvector, Redis, and OpenAI. Must have experience with: - Vector embeddings and semantic search - Document chunking and retrieval pipelines - Docker deployment on AWS
- →Stage A identifies archetype (AI/ML engineer), capabilities (RAG, vector search, LLM integration)
- →Stage B hits dense_hyde and capability channels for AI/ML projects in the corpus
- →Lexical channel catches exact tool names: FastAPI, pgvector, Redis
- →Proposal highlights past RAG, chatbot, or ML projects with architecture and scale details
Need a React Native developer for a cross-platform fitness tracking app. Features: GPS tracking, Apple Health / Google Fit integration, push notifications, offline mode. Budget: $5,000 fixed. Timeline: 8 weeks.
- →Stage A extracts mobile archetype, fitness domain, integration requirements
- →Graph retrieval finds projects near mobile + health capabilities even without exact keyword overlap
- →Judge scores trust and differentiator axes for portfolio relevance
- →Proposal emphasizes mobile delivery experience and relevant app store launches
Need a developer for a web project. Must know JavaScript. Budget negotiable.
- →Stage A may flag low signal — risk_flags in decomposition
- →CRAG (Stage E) may trigger if confidence is too low, broadening retrieval
- →Fewer high-confidence matches; proposal may note breadth of web/JS experience
- →Tip: paste more context (client history, skills list) for better results
API Reference
Base URL: http://127.0.0.1:8000
/healthzLiveness check — Postgres, Redis, vectors, Neo4j status
{ "postgres": true, "redis": true, "vectors": true, "neo4j": true, "degraded": [] }/healthz/criticalReadiness gate — returns 503 on embed-dim or corpus-version drift
{ "status": "ok" }/v1/proposals/streamStart full pipeline A→F; returns SSE event stream
{ "job_title": "...", "job_description": "...", "budget": "$2k-$3k", "skills": ["Python"], "options": { "skip_judge": false, "stop_after_match": false } }SSE events: stage.start → stage.complete (A,B,C,D,F) → proposal.delta → proposal.complete
Primary endpoint used by the UI. Use skip_judge: true for faster smoke tests.
/v1/proposals/{proposal_id}/streamRe-attach to an in-flight SSE run after disconnect
/v1/proposals/refineRe-run Stage F only with user instruction; returns SSE
{ "proposal_id": "...", "instruction": "Make it shorter and more casual", "pinned_project_ids": ["..."] }Reloads cached Stage A–D from job_sessions — no re-retrieval cost.
/v1/proposals/session/{session_id}/matchesRestore match-review state after SSE registry garbage collection
/v1/proposalsList proposals with pagination and filters
Query params: limit, offset, outcome (won|lost|pending|…), q (search text)
/v1/proposals/{proposal_id}Full proposal envelope: job, matched_projects, proposal, confidence, audit, outcome
/v1/proposals/{proposal_id}/feedbackRecord won/lost/interview outcome for offline correlation
{ "proposal_id": "...", "job_id": "...", "sent": true, "outcome": "interview", "user_rating": 4 }/v1/proposals/{proposal_id}/cancelCancel an in-flight run via Redis cancel flag
/v1/generateSync JSON API — query in, proposal + matches out (no SSE)
Useful for scripts and integrations; UI uses /stream instead.
/v1/settings/llmList LLM providers and active selection
/v1/settings/llm/testHealth-check a provider without activating it
/v1/settings/llm/activateValidate, persist, and hot-swap the active LLM provider
/v1/ingest/statusScheduler snapshot: enabled, interval, running, last_result, next_run_at
/v1/ingest/reportDated analytics: live pgvector + Neo4j sizes and ingest run history (datetime, new IDs, enrich/index/graph stats)
/v1/ingest/runFetch PCM and ingest only IDs missing from pgvector. 202 accepted, 409 if already running
Does not re-embed or re-graph projects that already have chunks.
Local Setup
Commands to bootstrap the full stack from scratch.
cd upwork-llm-api && docker compose up -dStart Postgres (:55432), Neo4j, Redis, MinIO
uv venv --python 3.11 --seed && source .venv/bin/activateCreate Python virtual environment
uv pip install -e ".[dev]"Install API dependencies
cp .env.example .envConfigure LLM_BASE_URL, DB ports to match compose.yaml
uv run alembic upgrade headApply database migrations
uv run upwork doctorVerify all backends are reachable
uv run upwork load-data data/raw/pcm_projects.jsonIngest raw project corpus
uv run upwork enrich --llmRun offline enrichment pipeline
uv run upwork indexBuild pgvector chunk index
uv run upwork build-graphBuild Neo4j capability graph
uv run upwork-apiStart FastAPI server (PCM ingest scheduler starts if PCM_SYNC_ENABLED=true)
cd ../upwork-llm-ui && npm install && npm run devStart Next.js UI on :3000