MI-RA Lab · Multimodal Intelligence for Real-World Applications

Egocentric Data Collection in India

What egocentric data is, how it differs from conventional video, and how researchers in India can plan first-person and embodied-AI capture in relation to MI-RA Lab’s controlled multimodal facility.

Search intent: researchers planning first-person, wearable, or embodied-AI data collection in India. Last updated 2026-09-12. Prepared by the MI-RA Lab team at IIT Mandi iHub and HCi Foundation. Technical claims on this page are limited to capabilities described on the lab overview and BookMyLab facility pages.

Direct answer

Egocentric data is first-person human data: video (and often audio, IMU, or gaze) from sensors on the head, chest, or wrists. It is the usual substrate for activity recognition from the wearer’s view, hand–object interaction, and many embodied or physical-AI pipelines. MI-RA Lab’s documented infrastructure is a controlled multi-view volumetric chamber, anechoic audio, OptiTrack motion capture, and multi-sensor rooms—not a published large-scale wearable-camera farm. Ego protocols can still be planned with the lab; kit and field-vs-chamber design must be confirmed, not assumed from this page.

Egocentric data, first-person video, and wearables

A head-worn camera sees what the person could see: hands entering the frame, objects on a table, other people in conversation. IMU tracks head or body acceleration and orientation. Gaze, when recorded, estimates where the eyes pointed. Audio from a wearable mic captures the wearer’s voice and nearby sound, with different acoustics from an anechoic boom.

Affiliate faculty pages on this site discuss research modalities such as eye movement, body motion, gestures, and IMU-class sensors. Those are research interests and possible protocol ingredients. They are not a warehouse inventory of issued ego glasses.

Hand–object interaction, embodied AI, robotics, and VLA models

Vision–language–action (VLA) and robot-learning work often needs to see hands, tools, and contact events from a viewpoint similar to a worker or a robot head. Egocentric video is closer to that viewpoint than a tripod in the corner. Third-person volumetric capture still helps: it can supply 3D body configuration and an external check on what the wearable missed.

Human–robot interaction studies may mix both: a participant wearing a camera while external cameras and mics record the scene. The multi-sensor room is described as flexible for UX and interaction layouts.

Conventional video versus egocentric data

Conventional third-person video compared with egocentric data
DimensionConventional / third-person videoEgocentric / first-person data
ViewpointExternal cameras; person is an object in the frameWearer-centred; hands and workspace dominate
Body appearanceFull body and face often visibleWearer’s face often absent; others appear from the front
Motion cuesOptical flow of the scene from a stable or tripod cameraStrong ego-motion; IMU often required to interpret shake
Best scientific useExpression, pose, multi-view reconstructionAttention, manipulation, navigation, embodied AI
Typical failure modeOcclusion of hands by the body; missed workspace detailMotion blur, privacy of bystanders, missing wearer face

Research-grade controlled capture versus large-scale field collection

Controlled lab capture compared with large-scale field egocentric collection
Controlled lab (e.g. MI-RA chambers)Large-scale field collection
StrengthRepeatable lighting and acoustics; easier ethics supervision; optional multi-view and physiologyNatural tasks, homes, workplaces, and Indian environments
LimitationTasks can look staged; wearers may behave formallySync, calibration, and bystander consent are harder
FitMethod papers, ablation of sensors, clean baselinesScale and ecological validity for activity models

Neither mode is “better.” Many strong studies use a controlled subset to debug alignment, then a field subset for validity. MI-RA is explicitly a controlled facility; field scale is a project choice, not a published lab product.

Limitations and evidence

  • No public ego dataset size, city coverage, or “largest in India” claim.
  • Wearable camera / eye-tracker models are not listed on the BookMyLab equipment cards the way the 16 Blackmagic cameras and OptiTrack are.
  • For third-person multi-view human data, use the volumetric capture description.

Direct answers

What is egocentric data collection?

Egocentric collection records the world from the participant’s own viewpoint—typically a wearable camera, often with IMU, audio, and sometimes gaze. It is first-person evidence of what the person attended to and manipulated, not a third-person stage recording.

Does MI-RA Lab host the largest egocentric dataset in India?

No such claim is made. This website does not publish egocentric hours, wearer counts, or a released ego corpus. Wearable-camera inventory is not listed as a fixed public facility specification.

What is the difference between volumetric capture and egocentric data?

Volumetric/multi-view capture uses many external cameras to observe the body in 3D. Egocentric data looks outward from the wearer. One reconstructs the person; the other reconstructs the wearer’s workspace and attention.

How can researchers collaborate with MI-RA Lab?

Describe whether you need third-person multi-view, first-person wearables, or both. Email mira@ihubiitmandi.in and use BookMyLab after project review.

Related MI-RA Lab pages

Collaborate with MI-RA Lab

Researchers and project teams can discuss capture protocols, ethics, and facility access with the lab. Booking is through BookMyLab after project review.

Email mira@ihubiitmandi.in · Contact and address · IIT Mandi iHub and HCi Foundation, North Campus, VPO Kamand, District Mandi, Himachal Pradesh, India - 175075