services/vision takes a real camera image and a list of text labels and returns named bounding boxes, so Scalpal can talk about what’s actually visible. It runs OWLv2 (Apache-2.0) on Apple Silicon with PyTorch MPS.
In the VR scene, Unity already knows every organ and tool exactly, so boxes there come from the scene graph. This service is only for real camera frames, like physical props and hands.
Run it
Limits
- About 0.75 s per frame on an M2, so roughly one frame a second. Fine for a coach, not for tracking.
- Similar instruments get confused. Use
candidateIdsand treat the label as a hint. - Boxes are 2D image coordinates with no depth.
- Measured on stock photos and synthetic renders only. No Quest passthrough frames have been tested.