Chapter 02 / Applied AI · Fraud signals
TrustKit AI.
Exploring fraud signals in rental property tours.
Source-verified processing budget
- Target frames per uploaded tour
- 7
- Quality measures · blur and brightness
- 2
- Tour frames used for listing comparison
- 3
- Entry paths · live and upload
- 2
The upload path targets seven sampled frames, bounding the number sent through frame analysis. For illustration, a 60-second, 30-fps video contains 1,800 frames: selecting seven sends about 99.6% fewer frames to that analysis stage.
Measurement scope & source
The reduction is arithmetic for the stated example, not a measured latency, cost or accuracy improvement. Short videos or decoding failures can yield fewer than seven frames. Optional listing comparison uses the first three frames; the combined assessment currently uses the first frame. Prototype processing counts do not establish fraud-detection accuracy.
The idea
Turning tour footage into useful observations—and questions worth asking before trusting a listing.
Python / FastAPI / OpenCV / Gemini / React / TypeScript
Animated architecture
From tour footage to reviewable signals
The map distinguishes live and uploaded-video paths. The deep-scan assessment currently uses combined first-frame observations.
01 / Capture
02 / Choose an entry path
03 / Extract two kinds of evidence
04 / Compare with claims
05 / Optional listing comparison
06 / Return for review
Control flow / Decisions & data
Two entry paths, different processing flows
Live Copilot and uploaded deep scans share analysis modules, but orchestrate them differently. The two lanes preserve those differences.
Scroll across the diagram to follow each branch
Reading the flow
- Listing claims can be enriched through a listing scrape when an address is provided. Live frame analysis runs the quality and vision modules concurrently.
- Uploaded video is sampled; quality and vision run sequentially for each extracted frame. The combined trust assessment currently uses the first frame’s observations.
- With an address, deep scan attempts a separate comparison using the first three tour frames and listing photos/details. Failure produces a comparison error entry.
- Vision mock fallback and live per-frame fallback paths exist. Assessments are investigation signals, not proof of fraud. Temporary upload cleanup runs in a finally block.
Why / What / How
The thinking behind the system.
Why a tour needs context
A remote renter sees a selective view of a property. A listing makes claims, while a tour provides incomplete visual evidence. TrustKit explores bringing those inputs together so a person can ask better follow-up questions before trusting what they see.
The useful output is an observation or inconsistency to investigate. Blur, darkness or an unusual frame can have ordinary explanations, so the product should preserve uncertainty rather than turn a visual heuristic into an accusation.
What I built within the team system
My work focused on frame extraction, metadata analysis and vision processing. Representative frames carry timestamps, and sampling limits and resizing bound the amount of media processed. The quality module measures blur using Laplacian variance and brightness using pixel intensity.
These measurements describe frame quality; they are not authentication of the video. Gemini vision adds observations, while the broader interface, voice experience, listing comparison and reasoning orchestration belong to the team project.
How live and uploaded tours differ
Live Copilot receives a JPEG frame over a WebSocket and runs quality and vision analysis concurrently. It combines those observations with listing claims, sends an alert response and continues the session. Per-frame failure has a fallback path so the connection can continue.
The uploaded-video path saves a temporary file, samples frames, then runs quality and vision in sequence for each frame. The current combined assessment uses the first frame’s observations even though the report contains multi-frame arrays. With an address, a separate comparison can use the first three tour frames and listing information. Temporary-file cleanup runs when processing ends.
Why those boundaries matter
Sampling keeps processing manageable but can miss a brief detail between frames. Using only the first frame for the combined assessment also limits how much of a tour that assessment represents. These are implementation constraints a reviewer should understand before interpreting the report.
Some vision paths return mock observations when configuration or inference fails. A successful interface response alone therefore does not demonstrate real model analysis or validated fraud detection. The next evaluation question is whether observations are useful and accurately tied to the media a renter actually supplied.
The starting point
Why this project?
Remote renters often have only a listing and a virtual tour to assess a property. The team explored whether voice, visual observations, and listing information could surface inconsistencies for further investigation.
My contribution
The work I brought to it.
My contribution focused on frame extraction, metadata analysis, and vision processing. The broader voice experience, listing comparison, reasoning pipeline, and interface are team work. This repository is a fork of the shared hackathon project.
How it works
From input to output.
A React and TypeScript interface connects to a FastAPI backend. OpenCV extracts representative frames with timestamps and sampling limits. The current metadata module measures frame blur and brightness, while a vision module integrates Gemini through Google Cloud. The team pipeline brings observations and listing information into an assessment.
- Tour media
- Sample frames
- Visual + listing signals
- Assessment for review
A design decision
Sample deliberately. Preserve the uncertainty.
Frame sampling, resizing, and a maximum frame count bound processing work instead of treating every video frame as equally useful. That keeps a prototype manageable but can miss brief details. Blur or low light can have ordinary causes; visual heuristics should prompt investigation, not become proof of fraud.
The evidence
Look under the surface.
These links point to the reviewed source revision, so the implementation behind this story stays inspectable.
Frame extraction
Inspect time-based sampling, frame limits, resizing, and per-frame timestamps.
backend/modules/frame_extractor.pyFrame quality analysis
The current implementation uses Laplacian variance and mean brightness, rather than the EXIF workflow in the early plan.
backend/modules/metadata_analyzer.pyVision integration
Inspect the cloud model connection and mock fallback when configuration or inference fails.
backend/modules/vision_analyzer.pyCurrent boundaries
Useful work. Honest limits.
- Hackathon prototype; no independently validated fraud-detection accuracy is claimed.
- Signals and assessments do not establish that a property listing is fraudulent.
- Cloud credentials and services are required for real model inference. Some vision paths fall back to mock observations.
- The early architecture plan describes capabilities beyond the current metadata implementation.
Source reviewed October 3, 2026.
Let’s build something
worth protecting.
“A thoughtful conversation can be the beginning of something worth building.”Connect on LinkedIn