Selected work

Chapter 02 / Applied AI · Fraud signals

TrustKit AI.

Exploring fraud signals in rental property tours.

Status

Team hackathon prototype

My role

Frame extraction, metadata analysis & vision processing

View source on GitHub

Source-verified processing budget

Target frames per uploaded tour
7
Quality measures · blur and brightness
2
Tour frames used for listing comparison
3
Entry paths · live and upload
2

The upload path targets seven sampled frames, bounding the number sent through frame analysis. For illustration, a 60-second, 30-fps video contains 1,800 frames: selecting seven sends about 99.6% fewer frames to that analysis stage.

Measurement scope & source

The reduction is arithmetic for the stated example, not a measured latency, cost or accuracy improvement. Short videos or decoding failures can yield fewer than seven frames. Optional listing comparison uses the first three frames; the combined assessment currently uses the first frame. Prototype processing counts do not establish fraud-detection accuracy.

Inspect frame sampling ↗Inspect the upload pipeline ↗Inspect quality measurements ↗

The idea

Turning tour footage into useful observations—and questions worth asking before trusting a listing.

Python / FastAPI / OpenCV / Gemini / React / TypeScript

Animated architecture

From tour footage to reviewable signals

The map distinguishes live and uploaded-video paths. The deep-scan assessment currently uses combined first-frame observations.

01 / Capture

02 / Choose an entry path

03 / Extract two kinds of evidence

04 / Compare with claims

05 / Optional listing comparison

06 / Return for review

Control flow / Decisions & data

Two entry paths, different processing flows

Live Copilot and uploaded deep scans share analysis modules, but orchestrate them differently. The two lanes preserve those differences.

Decision branchesSolid arrows show routing and return paths.

Scroll across the diagram to follow each branch

Two entry paths, different processing flowsLive Copilot and uploaded deep scans share analysis modules, but orchestrate them differently. The two lanes preserve those differences. A text explanation follows the diagram.No addressNext frameErrorContinueCompare if addressLive CopilotReact → WebSocket /ws/liveUploaded tourReact → POST /deep-scanReceive JPEG frameListing claims from session configSave & sample videoTemporary file → limited key framesAnalyze concurrentlyOpenCV quality + Gemini visionAnalyze each frameQuality then vision; repeat for framesMerge + reasonCurrent observations + listing claimsFirst-frame assessmentFirst quality + vision result + claimsSend alert JSONOptional warning audioAddress supplied?Optional listing/photo comparisonContinue sessionNext frame or contextual chatAssemble reportArrays + assessment + optional audioPer-frame failureFallback alert; keep session aliveReturn & clean upOptional comparison; delete temp file

Reading the flow

  1. Listing claims can be enriched through a listing scrape when an address is provided. Live frame analysis runs the quality and vision modules concurrently.
  2. Uploaded video is sampled; quality and vision run sequentially for each extracted frame. The combined trust assessment currently uses the first frame’s observations.
  3. With an address, deep scan attempts a separate comparison using the first three tour frames and listing photos/details. Failure produces a comparison error entry.
  4. Vision mock fallback and live per-frame fallback paths exist. Assessments are investigation signals, not proof of fraud. Temporary upload cleanup runs in a finally block.

Why / What / How

The thinking behind the system.

Why a tour needs context

A remote renter sees a selective view of a property. A listing makes claims, while a tour provides incomplete visual evidence. TrustKit explores bringing those inputs together so a person can ask better follow-up questions before trusting what they see.

The useful output is an observation or inconsistency to investigate. Blur, darkness or an unusual frame can have ordinary explanations, so the product should preserve uncertainty rather than turn a visual heuristic into an accusation.

What I built within the team system

My work focused on frame extraction, metadata analysis and vision processing. Representative frames carry timestamps, and sampling limits and resizing bound the amount of media processed. The quality module measures blur using Laplacian variance and brightness using pixel intensity.

These measurements describe frame quality; they are not authentication of the video. Gemini vision adds observations, while the broader interface, voice experience, listing comparison and reasoning orchestration belong to the team project.

How live and uploaded tours differ

Live Copilot receives a JPEG frame over a WebSocket and runs quality and vision analysis concurrently. It combines those observations with listing claims, sends an alert response and continues the session. Per-frame failure has a fallback path so the connection can continue.

The uploaded-video path saves a temporary file, samples frames, then runs quality and vision in sequence for each frame. The current combined assessment uses the first frame’s observations even though the report contains multi-frame arrays. With an address, a separate comparison can use the first three tour frames and listing information. Temporary-file cleanup runs when processing ends.

Why those boundaries matter

Sampling keeps processing manageable but can miss a brief detail between frames. Using only the first frame for the combined assessment also limits how much of a tour that assessment represents. These are implementation constraints a reviewer should understand before interpreting the report.

Some vision paths return mock observations when configuration or inference fails. A successful interface response alone therefore does not demonstrate real model analysis or validated fraud detection. The next evaluation question is whether observations are useful and accurately tied to the media a renter actually supplied.

The starting point

Why this project?

Remote renters often have only a listing and a virtual tour to assess a property. The team explored whether voice, visual observations, and listing information could surface inconsistencies for further investigation.

My contribution

The work I brought to it.

My contribution focused on frame extraction, metadata analysis, and vision processing. The broader voice experience, listing comparison, reasoning pipeline, and interface are team work. This repository is a fork of the shared hackathon project.

How it works

From input to output.

A React and TypeScript interface connects to a FastAPI backend. OpenCV extracts representative frames with timestamps and sampling limits. The current metadata module measures frame blur and brightness, while a vision module integrates Gemini through Google Cloud. The team pipeline brings observations and listing information into an assessment.

  1. Tour media
  2. Sample frames
  3. Visual + listing signals
  4. Assessment for review

A design decision

Sample deliberately. Preserve the uncertainty.

Frame sampling, resizing, and a maximum frame count bound processing work instead of treating every video frame as equally useful. That keeps a prototype manageable but can miss brief details. Blur or low light can have ordinary causes; visual heuristics should prompt investigation, not become proof of fraud.

The evidence

Look under the surface.

These links point to the reviewed source revision, so the implementation behind this story stays inspectable.

Current boundaries

Useful work. Honest limits.

  • Hackathon prototype; no independently validated fraud-detection accuracy is claimed.
  • Signals and assessments do not establish that a property listing is fraudulent.
  • Cloud credentials and services are required for real model inference. Some vision paths fall back to mock observations.
  • The early architecture plan describes capabilities beyond the current metadata implementation.

Source reviewed October 3, 2026.

Keep exploring / Chapter 03

SentinelScope

URL reputation checks through a familiar interface.

Let’s build something
worth protecting.

“A thoughtful conversation can be the beginning of something worth building.”
Connect on LinkedIn