Timely
Respond at the right moment—not merely as soon as possible.
ICASSP 2027 · SIGNAL PROCESSING GRAND CHALLENGE
HEARTBEAT evaluates whether human-centered systems can decide whether, when, and how to respond to conversational audio and continuous audio-visual or audio-only streams.
STATUS / REGISTRATION OPEN Registration has been open since October 7, 2026
Only information received up to the current time may be used.
Human-centered systems must distinguish genuine opportunities to respond from pauses, backchannels, self-repairs, ambiguous events, and moments where continued listening is better.
Knowing when not to act is part of acting intelligently.
Respond at the right moment—not merely as soon as possible.
Choose behavior that fits the conversation, current stream, and user goal.
Remain silent when the current evidence does not require a response.
Track I has separate Offline and Online evaluations and rankings. Track II offers Audio-Visual and Audio-Only subtracks, both evaluated through online streaming.
Know when to begin responding—and when to stop speaking.
Systems participate naturally in ongoing conversations by using the audio context and their current state to decide whether to respond, keep listening, stop speaking, or handle new information appropriately.
Requests · Clarification · Correction · Information updates · Interruptions · Pause, cancel, and resume
Natural pauses · Hesitations · Self-repairs · Incomplete utterances · Backchannels · Side speech
Decide when to respond and what to say as the stream unfolds.
Systems receive one instruction, then process audio-visual or audio-only input online and in chronological order. They must use only information observed so far, support multiple responses in one stream, and remain silent when no response is needed.
Current and previous chunks only · Continuous state · No access to the complete recording
Autonomous trigger detection · Multiple responses · Explicit silence when no response is needed
Online systems process each stream in chronological order and may access only the current chunk and information received earlier.
InputConversational audio prefix, assistant instructions, and four candidate actions
Required outputOne selected option per sample in JSON
InputReal-time chunks of user speech, other participants' speech, and environmental audio
Required outputOutput audio plus JSON logs with chunk intervals, speech onset and offset, and response content
InputAudio-visual or audio-only chunks received in chronological order
Required outputA JSON record for every chunk with its time interval and model output
Continuous state Systems maintain context as new input arrives. They cannot inspect the full recording in advance or defer every response until the stream ends.
Time-accounted output For Track II, a response is timestamped at the end of the chunk that produced it, so chunk length directly affects measured latency.
Every Track II chunk must be logged, including chunks with no response.
Each track follows its own task protocol. Correct content matters only when the system also responds at an appropriate time.
TOP-1 ACCURACY
Four-way selectionSelect one candidate action from the audio context at the designated decision point.
Exact scoringIncorrect, missing, or unparseable predictions are counted as errors.
No latency termOffline evaluation measures handling correctness only.
BEHAVIOR + QUALITY + TIMING
Turn-taking decisionCorrectly begin responding or stop speaking; missed, false, and duplicate triggers affect the result.
Response qualitySatisfy the request and context; different valid formulations are accepted.
Response timingMatch actual speech onset and cessation to their timing windows and report both latencies.
PRECISION / RECALL / JOINT F1
Temporal matchingMatch actual responses to ground-truth triggers inside the specified tolerance window.
Content correctnessVerify structured answers by task rules and open-ended responses against references.
Error sensitivityMissed events, false activations, and duplicate responses all reduce the score.
TRACK II TIMING A response is timestamped at the end of the chunk that generated it and cannot be backdated; chunk length therefore directly affects measured latency.
Public data and pretrained models are welcome when declared and legally usable. Award-eligible systems are evaluated under a transparent, frozen protocol.
Organizers and their PhD students are ineligible. Other institutional conflicts must be disclosed.
Public data and pretrained models must be declared, legally usable, and obtainable by the resource freeze.
Two development submissions per team/day; one primary and one backup final container.
Finalists provide a method card, manifests, model hashes, seeds, dependencies, and rerunnable code.
No-network OCI containers run on randomized filenames with state reset between episodes.
Private data, private test annotation, and private or commercial inference services.
All deadlines: 23:59 UTC unless otherwise stated. The final container deadline is highlighted.
Registration opens and website goes live
Data, rules, scoring tools, and baseline systems released
Development leaderboard opens
Team membership and resource declarations frozen
Development leaderboard closes
Final container submission deadline
Final rankings and paper invitations announced
Invited two-page papers due
Camera-ready two-page papers due
Up to five audited, top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person at the dedicated ICASSP 2027 SPGC session.
Invitations consider final rank, technical diversity, reproducibility, and balanced coverage of the two tracks.A cross-institutional team spanning audio research, benchmark design, data governance, infrastructure, and participant support.
Conference liaison, governance, conflict management, and paper and session coordination.
Dataset construction, annotation, privacy, task design, metrics, baselines, and hidden-test auditing.
Website, evaluation platform, containers, compute, participant support, webinars, and outreach.
Ant Group · National University of Singapore · The Chinese University of Hong Kong · University of Electronic Science and Technology of China · Institute of Computing Technology, Chinese Academy of Sciences · Singapore University of Technology and Design
No. Track I has separate Offline and Online evaluations and rankings. Track II has Audio-Visual and Audio-Only subtracks, and participants may register for either subtrack.
Track I uses conversational audio. Track II offers an audio-visual subtrack and an audio-only subtrack; participants may enter either one.
Systems process inputs in chronological order and may use only the current and previously received information. They cannot inspect the complete recording in advance or wait until the end to produce every answer.
Yes, if they are declared, legally usable, and obtainable by the November 20 resource-freeze date.
Track I Offline requires a JSON file with sample identifiers and selected options. Track I Online requires output audio and timestamped JSON logs. Track II requires per-chunk JSON containing time ranges and model outputs, including chunks with no response.
Track I Offline uses Top-1 accuracy; Track I Online evaluates turn-taking behavior, response quality, and response timing. Track II uses Precision, Recall, and joint F1 over correctly timed and correct responses.
Each track has a winner, and up to five audited top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person.
HEARTBEAT CHALLENGE 2027
Registration is open. Complete the official form to enter HEARTBEAT Challenge 2027. Data, rules, scoring tools, and baseline systems follow October 14; the development leaderboard opens October 21.
Register now