MI-RA Lab · Multimodal Intelligence for Real-World Applications

Synchronized Audio, Video and Biosignal Data Capture

Why files sitting in the same folder are not enough: temporal synchronization, clocks, latency, frame alignment, and metadata for audio, video, and biosignal capture at MI-RA Lab, IIT Mandi.

Search intent: engineers and researchers who need time-aligned audio, video, and biosignal recordings. Last updated 2026-09-12. Prepared by the MI-RA Lab team at IIT Mandi iHub and HCi Foundation. Technical claims on this page are limited to capabilities described on the lab overview and BookMyLab facility pages.

Direct answer

Having an MP4, a WAV, and a CSV from the “same session” is not enough. Those files are synchronized only if they can be mapped onto one timeline with known offset and drift. MI-RA Lab’s public description states that cameras, microphones, and other modalities are spatially arranged and synchronized using hardware-based techniques. This page explains what that claim must mean in practice. It does not publish a measured microsecond or millisecond accuracy figure, because none is given on the website.

Who, what, where, why, how

  • Who. Human-centric AI, HCI, speech, and physiology researchers using MI-RA Lab at IIT Mandi.
  • What. Time-aligned multi-view video, spatial audio, optional OptiTrack motion, and biosignals.
  • Where. Anechoic capture chamber and multi-sensor rooms; annotation and workstations for later QA.
  • Why. Cross-modal learning needs events, not independent clips.
  • How. Shared clocks or hardware triggers, calibration, session metadata—not folder names alone.

Temporal synchronization, timestamps, and trigger signals

Each device has a clock. Video cameras tick in frames (for example 24, 30, 60, or higher). Audio is sampled in kHz. EMG or EEG may run at hundreds to thousands of Hertz. If each device stamps only its local clock, the streams walk apart (clock drift) and may start at different wall-clock times (offset).

Common research practice—used across labs, and consistent with “hardware-based” synchronization—includes: a genlock or word-clock style reference; a shared SMPTE or PTP/NTP time where applicable; an electrical trigger that marks t = 0 on all recorders; or a clap / LED flash used only as a coarse check, not as the sole method for physiology.

Clock drift, sampling rates, sensor latency, and frame alignment

Drift is the slow change of one oscillator versus another. Latency is the fixed delay of a sensor or USB stack. Frame alignment is the mapping of video frame n to audio sample m and biosignal sample k. High-speed cameras and biosignals make sloppy alignment obvious: a 40 ms lip-sync error is visible; a 40 ms EMG-to-gesture error can invert a scientific conclusion.

Typical sampling regimes that must be mapped to one timeline
StreamTypical rate classAlignment question
Multi-view videoTens of frames per second (high-speed cameras higher)Are all 16 views genlocked or software-aligned?
Spatial audioTens of kHzIs there a common word clock or only file start time?
OptiTrack / mocapTens to hundreds of HzIs mocap time the master or a slave?
Biosignals (e.g. EMG/EDA)Often hundreds of Hz or moreWas the amp triggered or post-hoc stretched?

Exact rates for a booking are protocol-specific. Do not treat the table as MI-RA’s guaranteed configuration.

Calibration, metadata, session IDs, and quality control

  • Calibration. Camera extrinsics and mocap-to-camera transforms belong with the session, not in a forgotten spreadsheet.
  • Metadata. Session ID, operator, participant code, language, task, sensor serials, trigger method, and hash of raw files.
  • Quality control. The Annotation Room and workstations support review and labeling after capture.

Common timeline (conceptual)

The diagram is a teaching schematic. It is not a wiring diagram of the live rack.

Shared timeline for video, audio, and biosignalsA hardware clock and trigger feed video, audio, and biosignal recorders that write onto one session timeline with metadata.Reference clock+ trigger / genlock16-view videoSpatial audioBiosignalsCommon session timelineoffsets + drift documentedMetadata: session ID, calibration, sensor list, trigger methodWithout this block, folders are not a multimodal dataset
Conceptual sync path used to explain why hardware timing and metadata are part of capture, not post-production decoration.

Limitations

Do not cite this page as a datasheet. Ask the lab for the sync topology of your booking. Related reading: volumetric capture facility, Indian multimodal datasets, facility types compared.

Direct answers

What is synchronized multimodal data?

It is two or more sensor streams that share a common timeline so that an event at time t in video corresponds to the same t in audio, motion, and biosignals—within a stated error budget. Co-located files without shared time are not synchronized.

Why is temporal synchronization important?

Fusion models, lip-sync studies, and physiology-to-expression analyses fail or silently learn the wrong associations if one channel is shifted by tens of milliseconds. Hardware trigger and clock discipline exist to prevent that.

Where can I capture synchronized video and physiological signals?

MI-RA Lab’s capture and multi-sensor rooms are described as supporting simultaneous, hardware-synchronized video, spatial audio, and physiological data. Confirm the exact biosensor set and trigger method for your protocol.

Can audio, video and biosignals be collected simultaneously?

Yes, that is the stated purpose of the chamber and rig. Simultaneous start is still not the same as sample-accurate alignment; ask for the session’s sync method and metadata.

Related MI-RA Lab pages

Collaborate with MI-RA Lab

Researchers and project teams can discuss capture protocols, ethics, and facility access with the lab. Booking is through BookMyLab after project review.

Email mira@ihubiitmandi.in · Contact and address · IIT Mandi iHub and HCi Foundation, North Campus, VPO Kamand, District Mandi, Himachal Pradesh, India - 175075