Fine-tuned facebook/detr-resnet-50 on pixel-diff frame pairs; best val precision 0.60, recall 0.73.
2025
Tech stack
Problem
Fine-tuned the DETR (Detection Transformer) model on pixel-difference frame pairs to detect moved objects in video.
Approach
Ran systematic ablations on learning rate, backbone freezing, and augmentation strategies. Best checkpoint achieved val precision 0.60 and recall 0.73, with automated evaluation visualizations for every run..
Result
See the code for full results.
Full writeup
Fine-tuned the DETR (Detection Transformer) model on pixel-difference frame pairs to detect moved objects in video. Ran systematic ablations on learning rate, backbone freezing, and augmentation strategies. Best checkpoint achieved val precision 0.60 and recall 0.73, with automated evaluation visualizations for every run.