First-person action recognition and understanding from wearable cameras—where real life rarely holds still for the model.
where I got it wrong—
often enough to learn.
First-person action recognition and understanding from wearable cameras—where real life rarely holds still for the model.
Architectures that read motion across time, domains and viewpoints—not just the one clean clip used in the demo.
Teaching systems to survive a change of camera, factory or context without starting training from zero every time.
Finding rare, unusual or dangerous events when examples are scarce and “normal” keeps changing in production.
Detection, segmentation, recognition and pose estimation assembled into pipelines people can actually depend on.
Combining vision, audio, sensors and language when no single signal tells the whole story.
Stay close to people who can explain their no. That's how you build a yes that holds.
Surprise me.
The paper is
the clean version.
The useful part was messier:
wrong turns, rejected ideas & one more experiment.
Most of the learning never made the abstract.
mostly curiosity.
occasionally poor time management.
Papers, datasets and challenges for seeing the world from a first-person point of view.
Visit repository ↗
Tap a pixel, get the colour. HEX, RGB and CMYK—even offline.
Open project ↗
A quiet habit tracker built around consistency, not streak anxiety.
Open project ↗
Logos, roll-ups and visual experiments collected through the PhD years.
Browse the work →Good workdeserves agood presentation.
if you can't explain it simply,
the build isn't finished →
and a ruthless Q&A →
Unsupervised Domain Adaptation track — Top 3 for two consecutive years.
From hand-crafted descriptors to foundation models — one backbone to rule them all.
Presented cutting-edge first-person action recognition to Italy's Python AI community.
weird ideas welcome.