Timely
Respond at the right moment—not merely as soon as possible.
ICASSP 2027 · SIGNAL PROCESSING GRAND CHALLENGE
HEARTBEAT evaluates whether human-centered audio agents can decide whether, when, and how to respond to conversational and environmental audio.
STATUS / PRE-LAUNCH Website and registration open September 7, 2026
Appropriate abstention is a first-class outcome.
Human-centered audio agents must distinguish genuine opportunities to respond from pauses, backchannels, self-repairs, ambiguous events, and moments where continued listening or monitoring is better.
Knowing when not to act is part of acting intelligently.
Respond at the right moment—not merely as soon as possible.
Choose behavior that fits the conversation or acoustic scene.
Appropriate silence and continued monitoring are valid, rewarded outcomes.
The tracks share infrastructure and a calendar while retaining independent tasks, outputs, rankings, and winners.
Participate naturally in an ongoing conversation.
Given conversational audio context at designated decision points, systems decide whether to continue listening or take the turn. When taking the turn, they predict an appropriate response time and select an organizer-defined action.
Assistance requests · Clarification · Factual corrections · Safety cues
Pauses · Hesitations · Backchannels · Self-repairs · Incomplete utterances
Engage with an unfolding audio scene—only when it helps.
Given environmental or human non-speech audio, systems select silence, monitor, ask, or intervene. When action is warranted, they also provide a concrete response and a brief rationale.
Safety alarms · Door cues · Appliance states · Impacts · Human non-speech cues
Weak evidence · Premature moments · Resolved events · No-action scenes
All inputs are 10–90 second audio segments sampled at 16 kHz, with reproducible tooling from training to hidden evaluation.
Hidden Main follows the target-condition distribution while remaining disjoint in speakers, sessions, and source recordings.
Hidden Generalization introduces predefined shifts in speakers, environments, devices, sound sources, interaction patterns, and acoustic conditions.
Private evaluation · Submitted OCI containers · Frozen scorers · No network access
Each component is normalized to [0, 1]. A single score is computed for each track over the combined hidden benchmark.
0.30 Decision + 0.35 Timing + 0.35 Action
Turn-taking decisionF1 for whether to take the turn.
Response timingAccuracy within temporal tolerance.
Action selectionTop-1 accuracy for the correct action.
0.30 Timing + 0.30 Action + 0.40 Response
Engagement timingAccuracy of an appropriate engagement time.
Action predictionTop-1 accuracy over four actions.
Response + rationaleCorrectness and appropriateness.
TIE-BREAK Hidden Generalization performance, then lower response latency for Track I or lower false-activation rate for Track II.
Public data and pretrained models are welcome when declared and legally usable. Award-eligible systems are evaluated under a transparent, frozen protocol.
Organizers and their PhD students are ineligible. Other institutional conflicts must be disclosed.
Public data and pretrained models must be declared, legally usable, and obtainable by the resource freeze.
Two development submissions per team/day; one primary and one backup final container.
Finalists provide a method card, manifests, model hashes, seeds, dependencies, and rerunnable code.
No-network OCI containers run on randomized filenames with state reset between episodes.
Private data, private test annotation, and private or commercial inference services.
All deadlines are at 23:59 UTC unless stated otherwise. The final container deadline is highlighted.
Registration and website open
Data, rules, scorers, and baselines released
Development leaderboard opens
Teams and resource declarations freeze
Development leaderboard freezes
Final container submission deadline
Rankings and paper invitations announced
Invited 2-page papers due
Camera-ready 2-page papers due
Up to five audited, top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person at the dedicated ICASSP 2027 SPGC session.
Invitations consider final rank, technical diversity, reproducibility, and balanced coverage of the two tracks.A cross-institutional team spanning audio research, benchmark design, data governance, infrastructure, and participant support.
Conference liaison, governance, conflict management, and paper and session coordination.
Dataset construction, annotation, privacy, task design, metrics, baselines, and hidden-test auditing.
Website, evaluation platform, containers, compute, participant support, webinars, and outreach.
Ant Group · National University of Singapore · The Chinese University of Hong Kong · University of Electronic Science and Technology of China · Institute of Computing Technology, Chinese Academy of Sciences · Singapore University of Technology and Design
No. They share infrastructure and a schedule but have independent tasks, outputs, rankings, and winners.
Both tracks use 10–90 second segments sampled at 16 kHz. Track I focuses on conversational speech; Track II focuses on environmental and human non-speech audio.
Yes, if they are declared, legally usable, and obtainable by the November 20 resource-freeze date.
Organizers execute submitted no-network OCI containers against private Hidden Main and Hidden Generalization benchmarks using frozen scorers.
Each track has a winner, and up to five audited top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person.
HEARTBEAT CHALLENGE 2027
Registration, documentation, leaderboard access, tutorials, webinars, and the participant forum launch September 7, 2026.
Save the dates