← Back to work

BidVisory

Property owners and managers compare inconsistent contractor bids without a repeatable way to defend the decision.

A source-aware evaluation workflow that extracts competing proposals, verifies arithmetic, researches vendors and market context, and composes an owner-ready report.

My role · Solo founder

  • Problem research and product strategy
  • Product and report design
  • Full-stack and AI workflow engineering
  • Deployment, privacy, cost, and operations
Real product walkthrough1:15 · 16:9
Read transcript
  1. 01BidVisory starts with the source proposals, including text PDFs, Word documents, scans, and phone photos.
  2. 02Parallel extraction normalizes scope, prices, dates, terms, and vendor identity.
  3. 03Research is restricted to eligible source URLs while arithmetic and scoring rules stay in code.
  4. 04The result is a navigable, owner-ready comparison with red flags and questions to ask.
  5. 05Cost, latency, structured outputs, source URLs, and evidence links are monitored continuously; field accuracy still requires labeled review data.
Structured bids
25

Retained validator-conforming records

Mean vendors
2.24

Per report

Expected report cost
$1.35

Current rates and observed mix

Text-PDF extraction
$0.009

One-page measured document

01 / Problem and workflow

From proposals to a defensible decision

BidVisory was designed around the decision artifact, not a chat box. The report must show its work, expose gaps, and survive scrutiny.

Before

Collect PDFs, Word files, scans, and scope documents from vendors.

Read line items and normalize scope manually across inconsistent proposal formats.

Search vendors, licenses, reputation, references, materials, and local market pricing.

Draft a side-by-side comparison, recommendation, red flags, and follow-up questions.

After

Upload the proposal pack without filling a long intake form.

Inspect structured extractions and vendor identities before research is composed.

Receive one consistent report with normalized scope, sourced context, and deterministic math.

Use the report as an owner or board packet instead of rebuilding a spreadsheet.

Four-stage durable workflow

  1. 01

    Extract

    Parse each proposal independently into structured scope, price, terms, timeline, warranty, and vendor fields.

  2. 02

    Identify

    Resolve vendor identity from the bid and Places data before web research begins.

  3. 03

    Research

    Collect market and reputation evidence with an eligible-URL set retained for downstream citations.

  4. 04

    Compose

    Generate a normalized report after deterministic arithmetic, scoring, and integrity checks.

Code-owned arithmetic

Totals, averages, adjustments, and reconciliation are computed and tested outside the model.

Fixed scoring rubric

Criteria live in versioned application code so the evaluation posture is explicit and repeatable.

Source URL allowlist

Research claims may cite only URLs gathered in the research stage; unknown URLs are removed.

Retries and fallbacks

Transient failures retry at the step boundary; exhausted optional research degrades explicitly.

02 / Shipped report

The output is the interface

The report keeps the recommendation, comparison, reputation evidence, market context, warnings, and vendor questions in one navigable document.

BidVisory public sample report with section navigation and summary

No-login sample report

Hiring managers can inspect the actual report anatomy without creating an account. Sample vendors and prices are clearly labeled illustrative.

BidVisory sample report focused on its recommendation and bid comparison

Recommendation hierarchy

The report leads with a recommendation and rationale, then moves into comparable bids, tradeoffs, warnings, and follow-up questions.

Architecture choices

Extraction runs in parallel by document; deterministic normalization and arithmetic happen before report composition.

Vendor identity combines proposal evidence and Places data. Research returns an explicit eligible URL set used by the composition prompt and post-generation validation.

Durable workflow steps retry independently. Completed journals are cleaned because they contain extracted customer document text.

  1. 01

    Documents

    Text PDFs, scans/photos, and DOCX proposals.

  2. 02

    Parallel extraction

    Independent parsing avoids one bad file killing the entire pack.

  3. 03

    Normalize and verify

    Units, totals, dates, and arithmetic are checked in code.

  4. 04

    Identity and research

    Places plus allowlisted web sources establish context.

  5. 05

    Compose and validate

    The report is assembled, scrubbed, and stored as the durable artifact.

  6. 06

    Privacy cleanup

    Completed workflow journals are removed after the durable report is stored.

03 / Evaluation

Measured & Monitored

Recorded usage supports the cost model. Operational telemetry is captured continuously, and the workflow is reviewed and adjusted when models, prompts, pricing, or behavior change.

Measured unit economics

Mean vendors per report

17-report production snapshot

2.24

Text-PDF extraction

Per one-page measured document

$0.009

Three-vendor fixture

Nine AI calls + three Places calls

$1.56

Expected report cost

Modeled from the current vendor mix

$1.35

90th percentile report cost

90% of modeled reports cost this or less

$1.74

Reputation research

Share of measured three-vendor run

59%

Download sanitized evidence JSON ↓

Monitored

These signals describe completion, latency, and evidence controls. They do not represent semantic extraction accuracy.

Report latency

90% of 17 retained reports finalized within this time

2m 43s

Structured bid records

Validator-conforming extractions retained; 3 drafts excluded

25

Stored source URLs

HTTP(S) URL present; not claim-level citation coverage

1,321 / 1,321

Rating evidence links

Stored rating entries with an evidence or profile link

57 / 57

Pipeline guardrail tests

Identity, provenance, arithmetic, scoring, and fail-closed behavior

127 passing

Field-level extraction precision and recall, unsupported-claim rate, expert agreement, and human correction rate still require labeled review data before publication.

Decisions and rejected alternatives

Starter allowance moved from 10 to 8 reports and Pro from 40 to 25 after measured cost logs exposed the full-utilization tail.

Vendor-count metering was rejected because the customer does not control how many vendors bid the job.

Reputation caching was rejected by owner decision; freshness and fairness outweighed the largest theoretical saving.

Workflow journals are cleaned after completion because intermediate payloads contain customer document text.

What the sample changed

In the July usage fixture, only 4% of 25 extracted bids printed a vendor website and 8% printed a phone number.

Places matched 5 of 35 vendor observations in that fixture, causing the more expensive profile-discovery fallback more often than assumed.

Applying current rates to the September report mix produces a model of approximately $0.21 plus $0.51 per vendor: $1.35 expected, with 90% of modeled reports costing $1.74 or less.

The next quality claim requires a scored benchmark; this portfolio does not substitute architecture language for that evidence.

Next case study

LineStriper AI

Parking-lot striping contractors lose time moving from aerial takeoff to estimate to signed quote.

Read case study →