Run your own video
Start with the included Colab notebook or install locally with Python and FFmpeg. Generated images and enrichment use optional paid model stages.
Open the notebook in Colab ↗ClosetScan extracts garments from a phone video, groups their views, and adds generated flat-lays, attributes, and narration. Explore a real run below.
Video → candidate garment frames
Distinct pieces, front/back images, attributes
Interactive catalogue + local MCP server
Precomputed results · No API key needed to explore
Select a garment, then a timestamp to jump to its source in the recording.
These results were generated from this recording using DINOv2 and the paid AI stages.
Ordered by first appearance in the video.
Flat-lays are AI-generated reconstructions and may invent details. Open a garment to inspect its attributes, narration, and original frames.
Process your own recording, inspect the outputs, or give an agent access to the catalogue.
Start with the included Colab notebook or install locally with Python and FFmpeg. Generated images and enrichment use optional paid model stages.
Open the notebook in Colab ↗The local MCP server exposes garment search, details, narration, and wear logs. The included wardrobe skill shows agents how to use it.
View MCP configuration ↗Read the setup instructions, pipeline stages, and known limits. ClosetScan is MIT licensed; contributions are welcome.
Read the project README ↗