Dean Shabi
All work

Katalo · 2026

AI can stage a living room. It shouldn't move the walls.

Katalo staged, renovated and decluttered listing photos for real-estate agencies. Image models are good at furniture and bad at walls, and a photo that misrepresents a property can't be published. I built the pipeline that generated each edit, checked it and repaired the ones that failed.

Role
Co-founder · AI pipeline, evals and infrastructure
Stack
TypeScript, Convex, Gemini, FAL, OpenRouter
A furnished living room with a sofa, an armchair, a stone coffee table and a glass wall onto a garden

How it works

  1. 1

    Listing photo

  2. 2

    Image generation

  3. 3

    Quality judge

  4. 4

    Repair or approve

  5. 5

    Delivery

A vision model checked every edit against the rules human editors followed. A failed edit came back with fix instructions for the next attempt.

Results

How the judge decided

A vision model compared every edited photo with the original and scored it against the rubric human editors used. The model only reported. Plain code made the call. Of the edits it approved, 95% were approved by human reviewers too.

The judge saw

  • Original photo
  • Edited photo
  • Editors' rubric

It returned

  • Score from 1 to 5
  • Structural failures
  • Fix instructions

Code decided

Score of 4 or more and no structural failure

  • Yes. Publish to the listing
  • No. Retry with the fix instructions, up to 3 attempts

Precision measured against human reviewers.

Agencies could also call the pipeline through an API. A retried request never generated twice, and completion webhooks were signed.

Key decisions

  1. 01

    Turn the editing handbook into a rubric

    I rewrote the handbook human editors used as a structured rubric for a vision model. The model scored each edit and flagged structural failures like a moved window. Plain code then decided whether the image could ship.

  2. 02

    Use rejections as repair instructions

    When the judge rejected an edit, it said what to fix. Those instructions went into the next attempt, which could also switch to a different model family. Attempts were capped.

  3. 03

    Measure what the agency would see

    I calibrated the judge against human labels and replayed each listing to see which image would actually have been published. Wrong approvals and wrong rejections were counted separately, because they cost different things.

  4. 04

    Share four providers without falling over

    Each provider had its own rate limits and failure modes. Per-customer limits stopped one bulk upload from blocking everyone else, and the system slowed down on its own when a provider pushed back.

Have a machine learning system that has to hold up in production? I'd like to hear about it.