Skip to main content
Upwork Proposal Intelligence

Upwork Proposal Intelligence

Paste a job — read a grounded proposal.

Overview

Upwork Proposal Intelligence helps BD and freelancer operators write stronger Upwork proposals by finding the most relevant past company projects (~139 PCM projects) and generating a citation-grounded proposal that cites only real work and real URLs.

The system is split into two apps: a Python FastAPI backend (intelligence engine) and a Next.js frontend (operator workspace). They communicate over REST and Server-Sent Events (SSE) for live pipeline progress.

The problem: Raw project descriptions are messy prose with noisy tags (REACT vs REACT.JS). Simple keyword or single-vector search cannot pick the right past work. This is a structured-knowledge-from-unstructured-prose problem — solved by enriching projects offline and running a multi-stage retrieval + judge pipeline online.

System architecture

UI
Next.js workspace on :3000 — paste job, watch live stages, review matches, refine, track history
API
FastAPI on :8000 — decompose, hybrid retrieve, rerank, LLM judge, write proposal
Postgres + pgvector
Projects, sessions, proposals, and dense chunk embeddings
Neo4j
Capability graph for proximity-based retrieval
Redis
Cache embeddings & decompositions; cancel in-flight runs
LLM / Embed / Rerank
Ollama, OpenRouter, Groq, or Gemini via OpenAI-compatible clients

Technology Stack

Each technology was chosen for a specific reason. Below is what we use, and why.

Backend

ComponentTechnologyWhy we use it
LanguagePython ≥ 3.11Strong typing, async I/O, mature LLM/ML ecosystem
Packaginguv + hatchlingFast installs; CLI entrypoints upwork and upwork-api
APIFastAPI + uvicornAsync HTTP, clean OpenAPI, streaming-friendly
Streamingsse-starlettePush live stage events (decompose → retrieve → write) to the UI
Configpydantic-settingsTyped .env config for all backends
ORM / MigrationsSQLAlchemy async + AlembicEvolve Postgres schema safely
Observabilitystructlog, Prometheus, OpenTelemetryPer-stage latency, tokens, correlation IDs

Data stores

ComponentTechnologyWhy we use it
Primary DB + vectorsPostgres 16 + pgvectorOne source of truth for projects, sessions, proposals, and dense embeddings
Full-text searchPostgres FTS / BM25Exact tech/tool name hits that embeddings can miss
Capability graphNeo4j 5Projects near these capabilities — proximity beyond keyword overlap
Cache / cancelRedisCache embeddings & job decompositions; cancel in-flight SSE runs
Object storeMinIO (compose)Optional artifacts storage for local/demo stacks

AI / ML models

ComponentTechnologyWhy we use it
Heavy LLMgpt-oss:120b via Ollama (default)Decompose job, judge matches, write proposal — needs strong reasoning
Cheap LLMqwen2.5:7bLighter helper calls where full heavy model is unnecessary
EmbeddingsQwen3-Embedding-8B (VECTOR_DIM=1024)Dense similarity for HyDE paragraphs and sub-questions
RerankerBGE-reranker (TEI or in-process)Cross-encode job vs candidates; precision after fusion
Provider swapOpenAI-compatible clientsChange base URL / profile to switch Ollama ↔ cloud without rewriting pipeline

Frontend

ComponentTechnologyWhy we use it
FrameworkNext.js 14 (App Router) + React 18Fast SPA-style workspace for operators
StylingTailwind + shadcn / Base UIConsistent, accessible controls
Server stateTanStack React QueryHistory and proposal envelope fetches
Live updates@microsoft/fetch-event-sourceReliable SSE with reconnect for pipeline theatre
URL statenuqsDeep-link ?proposal= and ?review= for sharing and reload
TestsVitestUnit coverage for stream/API helpers

Pipeline Steps

The system has two flows: an offline flow to prepare the corpus (run once or on updates), and an online flow for every job submission.

Offline flow — prepare the corpus

01Load Dataupwork load-data

Why: Persist the truth source — names, descriptions, tags, URLs from pcm_projects.json

Postgres projects table (raw rows)

02Enrichupwork enrich --llm

Why: Fix tag noise; add domains, capabilities, stack, architecture, scale, outcomes via LLM

ProjectIntelligence JSONB in Postgres

03Indexupwork index

Why: Enable dense semantic retrieval over project chunks

pgvector HNSW index on chunks.embedding

04Build Graphupwork build-graph

Why: Enable capability-proximity retrieval beyond keyword overlap

Neo4j capability graph from ontology YAMLs

05Daily ingest (new IDs only)scheduler + POST /v1/ingest/run

Why: Keep knowledge current when PCM adds projects — without re-formatting or re-embedding the old corpus

New IDs only: LLM format → pgvector chunks → Neo4j Project nodes

Online flow — every job submission

AStage A — DecomposeLLM understands job

Why: Aligns the job to the same schema as enriched projects; HyDE paragraphs and sub-questions improve recall

DecomposedJob: domains, capabilities, HyDE paras, sub-questions, archetype. Cached in Redis (~24h).

BStage B — Hybrid Retrieve5 parallel channels + RRF fusion

Why: No single channel is enough — semantic, structured, graph, and lexical each catch different misses

~top 25 candidates fused via Reciprocal Rank Fusion (k=60): dense_hyde, dense_subq, capability filter, Neo4j graph, Postgres FTS

CStage C — RerankCross-encoder precision

Why: Cheap fusion is recall-heavy; cross-encoder adds precision before the expensive judge

BGE reranker scores job vs candidates → ~top 12 survivors

DStage D — LLM Judge7-axis rubric scoring

Why: Human-readable fit scores + rationales; quality floor for what the writer may cite

Ranked matches with domain, stack, capability, architecture, scale, trust, differentiator scores

EStage E — CRAG (conditional)Corrective retrieval retry

Why: Self-check when retrieval looks weak; retry smarter before writing garbage

Broadened retrieval if confidence too low, then re-run judge

FStage F — WriteGrounded proposal generation

Why: Produce the actual proposal; linters block invented project IDs, URLs, and capabilities

Citation-grounded proposal text streamed via SSE proposal.delta events

Refine flow: POST /v1/proposals/refine reloads cached Stage A–D and re-runs only Stage F with your instruction. Editing tone should not re-pay retrieve + judge cost.

Daily knowledge ingest (new projects only)

When PCM publishes a project that is not already in the vector store, the API formats it, embeds it into Postgres pgvector, and adds it to the Neo4j graph. Projects that already have rows in chunks are skipped — no re-format, no re-embed, no graph rebuild.

01Fetch PCM

Download the latest project list and merge into data/raw/pcm_projects.json.

02Detect new IDs

New = PCM / corpus ID with zero rows in Postgres chunks. Already-indexed IDs are never reprocessed.

03Format (LLM)

Enrich only those IDs if enriched is still null.

04Embed → pgvector

Chunk + embed only those IDs into the chunks table (this is the vector store — there is no separate Qdrant).

05Graph → Neo4j

Upsert Project subgraphs only for those IDs.

How to verify vector + graph

pgvector (Postgres chunks)
SELECT COUNT(DISTINCT project_id), COUNT(*) FROM chunks; new IDs must appear in DISTINCT project_id.
Neo4j graph
MATCH (p:Project) RETURN count(p); new IDs exist as Project nodes after graph.done.
API report
GET /v1/ingest/report (this page, Ingest analytics) — live counts plus dated run history.

Scheduler env

VariableTypicalMeaning
PCM_SYNC_ENABLEDtrueStart the background scheduler with the API
PCM_SYNC_INTERVAL_HOURS24Hours between ticks. Floor is 60 seconds (use 0.0167 for a 1-minute test)
PCM_SYNC_USE_LLMtrueRun LLM format on new IDs
PCM_SYNC_RUN_ON_STARTUPfalseIf true, run once immediately when the API starts

Ingest analytics

Live pgvector and Neo4j sizes, plus dated ingest runs (when new IDs were actually written). History is saved under data/ingest_reports.json whenever a run syncs or errors — skipped ticks are not logged.

Snapshot time: —

Loading live store counts…

How to Use

Step-by-step guide for operators using the workspace UI.

  1. 1
    Open the workspace

    Action: Navigate to /workspace (or click Open Workspace from the landing page).

    Result: You see the composer and an empty state with a sample job option.

  2. 2
    Paste the Upwork job

    Action: Copy the full job description from Upwork — title, requirements, budget, skills. Paste into the composer.

    Result: The job appears as a user message in the thread.

  3. 3
    Submit and watch the pipeline

    Action: Click Submit. The pipeline theatre shows live progress: Decompose → Retrieve → Rerank → Judge → Write.

    Result: SSE events stream stage completions and proposal text tokens in real time.

  4. 4
    Review matched projects (optional)

    Action: If match review is enabled (stop_after_match), pin or drop projects before generation continues.

    Result: Only selected projects are cited in the final proposal.

  5. 5
    Read the proposal

    Action: Scroll through the generated proposal. Expand the audit drawer to see matched projects, scores, and confidence.

    Result: A grounded proposal citing real past projects with real URLs — no invented work.

  6. 6
    Refine if needed

    Action: Use the refine panel to adjust tone, length, or focus. Refinement re-runs only Stage F using cached matches.

    Result: An updated proposal appended as a new turn — faster than a full re-run.

  7. 7
    Record outcome

    Action: Mark the proposal as won, lost, interview, no reply, etc. from the outcome modal.

    Result: Feedback stored for offline correlation — helps improve future ranking signals.

  8. 8
    Browse history

    Action: Use the sidebar to filter by outcome and reopen past proposals.

    Result: Full envelope restored: job, matches, proposal text, audit trail.

Query Examples

What to paste in the composer, and what the system returns for each type of job.

Full-stack fintech dashboard
Input (paste this)
We're a Series A fintech startup looking for a senior full-stack engineer to build a real-time payments dashboard.

Requirements:
- React/Next.js front end with live-updating charts over WebSockets
- Node.js + PostgreSQL backend, Stripe Connect integration
- Handle 10k+ transactions/day at sub-second latency
- SOC 2 awareness a plus

3-month contract with potential to extend.
What you get
  • →Stage A extracts domains (fintech, payments), capabilities (real-time dashboards, Stripe integration), and tech stack signals
  • →Stage B retrieves past projects with React dashboards, PostgreSQL backends, payment integrations, and WebSocket experience
  • →Stage D scores matches on domain fit, stack overlap, and scale handling
  • →Stage F writes a proposal citing 5–10 relevant past projects with specific URLs and outcomes
AI / RAG chatbot engineer
Input (paste this)
Looking for an experienced Python developer to build a production-ready RAG chatbot using FastAPI, PostgreSQL, pgvector, Redis, and OpenAI.

Must have experience with:
- Vector embeddings and semantic search
- Document chunking and retrieval pipelines
- Docker deployment on AWS
What you get
  • →Stage A identifies archetype (AI/ML engineer), capabilities (RAG, vector search, LLM integration)
  • →Stage B hits dense_hyde and capability channels for AI/ML projects in the corpus
  • →Lexical channel catches exact tool names: FastAPI, pgvector, Redis
  • →Proposal highlights past RAG, chatbot, or ML projects with architecture and scale details
Mobile app (React Native)
Input (paste this)
Need a React Native developer for a cross-platform fitness tracking app.

Features: GPS tracking, Apple Health / Google Fit integration, push notifications, offline mode.
Budget: $5,000 fixed. Timeline: 8 weeks.
What you get
  • →Stage A extracts mobile archetype, fitness domain, integration requirements
  • →Graph retrieval finds projects near mobile + health capabilities even without exact keyword overlap
  • →Judge scores trust and differentiator axes for portfolio relevance
  • →Proposal emphasizes mobile delivery experience and relevant app store launches
Vague / short job post
Input (paste this)
Need a developer for a web project. Must know JavaScript. Budget negotiable.
What you get
  • →Stage A may flag low signal — risk_flags in decomposition
  • →CRAG (Stage E) may trigger if confidence is too low, broadening retrieval
  • →Fewer high-confidence matches; proposal may note breadth of web/JS experience
  • →Tip: paste more context (client history, skills list) for better results

API Reference

Base URL: http://127.0.0.1:8000

Open interactive Swagger docs
GET/healthz

Liveness check — Postgres, Redis, vectors, Neo4j status

{ "postgres": true, "redis": true, "vectors": true, "neo4j": true, "degraded": [] }
GET/healthz/critical

Readiness gate — returns 503 on embed-dim or corpus-version drift

{ "status": "ok" }
POST/v1/proposals/stream

Start full pipeline A→F; returns SSE event stream

{ "job_title": "...", "job_description": "...", "budget": "$2k-$3k", "skills": ["Python"], "options": { "skip_judge": false, "stop_after_match": false } }
SSE events: stage.start → stage.complete (A,B,C,D,F) → proposal.delta → proposal.complete

Primary endpoint used by the UI. Use skip_judge: true for faster smoke tests.

GET/v1/proposals/{proposal_id}/stream

Re-attach to an in-flight SSE run after disconnect

POST/v1/proposals/refine

Re-run Stage F only with user instruction; returns SSE

{ "proposal_id": "...", "instruction": "Make it shorter and more casual", "pinned_project_ids": ["..."] }

Reloads cached Stage A–D from job_sessions — no re-retrieval cost.

GET/v1/proposals/session/{session_id}/matches

Restore match-review state after SSE registry garbage collection

GET/v1/proposals

List proposals with pagination and filters

Query params: limit, offset, outcome (won|lost|pending|…), q (search text)

GET/v1/proposals/{proposal_id}

Full proposal envelope: job, matched_projects, proposal, confidence, audit, outcome

POST/v1/proposals/{proposal_id}/feedback

Record won/lost/interview outcome for offline correlation

{ "proposal_id": "...", "job_id": "...", "sent": true, "outcome": "interview", "user_rating": 4 }
POST/v1/proposals/{proposal_id}/cancel

Cancel an in-flight run via Redis cancel flag

POST/v1/generate

Sync JSON API — query in, proposal + matches out (no SSE)

Useful for scripts and integrations; UI uses /stream instead.

GET/v1/settings/llm

List LLM providers and active selection

POST/v1/settings/llm/test

Health-check a provider without activating it

POST/v1/settings/llm/activate

Validate, persist, and hot-swap the active LLM provider

GET/v1/ingest/status

Scheduler snapshot: enabled, interval, running, last_result, next_run_at

GET/v1/ingest/report

Dated analytics: live pgvector + Neo4j sizes and ingest run history (datetime, new IDs, enrich/index/graph stats)

POST/v1/ingest/run

Fetch PCM and ingest only IDs missing from pgvector. 202 accepted, 409 if already running

Does not re-embed or re-graph projects that already have chunks.

Local Setup

Commands to bootstrap the full stack from scratch.

1.
cd upwork-llm-api && docker compose up -d

Start Postgres (:55432), Neo4j, Redis, MinIO

2.
uv venv --python 3.11 --seed && source .venv/bin/activate

Create Python virtual environment

3.
uv pip install -e ".[dev]"

Install API dependencies

4.
cp .env.example .env

Configure LLM_BASE_URL, DB ports to match compose.yaml

5.
uv run alembic upgrade head

Apply database migrations

6.
uv run upwork doctor

Verify all backends are reachable

7.
uv run upwork load-data data/raw/pcm_projects.json

Ingest raw project corpus

8.
uv run upwork enrich --llm

Run offline enrichment pipeline

9.
uv run upwork index

Build pgvector chunk index

10.
uv run upwork build-graph

Build Neo4j capability graph

11.
uv run upwork-api

Start FastAPI server (PCM ingest scheduler starts if PCM_SYNC_ENABLED=true)

12.
cd ../upwork-llm-ui && npm install && npm run dev

Start Next.js UI on :3000