ICASSP 2027 · SIGNAL PROCESSING GRAND CHALLENGE

Teach audio agents when to act— and when not to.

HEARTBEAT evaluates whether human-centered audio agents can decide whether, when, and how to respond to conversational and environmental audio.

STATUS / PRE-LAUNCH Website and registration open September 7, 2026

LIVE AUDIO CONTEXT00:42.8
DECISION WINDOWContinue listening
92.4%
SILENCEMONITORASKINTERVENE

Appropriate abstention is a first-class outcome.

02Independent tracks
10–90sAudio segments
16 kHzSampling rate
70 / 30Main / generalization
01 / THE QUESTION

Listening is not enough. An agent must know what to do next.

Human-centered audio agents must distinguish genuine opportunities to respond from pauses, backchannels, self-repairs, ambiguous events, and moments where continued listening or monitoring is better.

Knowing when not to act is part of acting intelligently.

01

Timely

Respond at the right moment—not merely as soon as possible.

02

Context-aware

Choose behavior that fits the conversation or acoustic scene.

03

Restrained

Appropriate silence and continued monitoring are valid, rewarded outcomes.

02 / THE TRACKS

Two tracks.
Two independent rankings.

The tracks share infrastructure and a calendar while retaining independent tasks, outputs, rankings, and winners.

TRACK I01

Reciprocal Audio Turn-Taking

Participate naturally in an ongoing conversation.

Given conversational audio context at designated decision points, systems decide whether to continue listening or take the turn. When taking the turn, they predict an appropriate response time and select an organizer-defined action.

  1. 01Turn-taking decision
  2. 02Response timing
  3. 03Action selection
POSITIVE CASES

Assistance requests · Clarification · Factual corrections · Safety cues

HARD NEGATIVES

Pauses · Hesitations · Backchannels · Self-repairs · Incomplete utterances

TRACK II02

Proactive Audio Engagement

Engage with an unfolding audio scene—only when it helps.

Given environmental or human non-speech audio, systems select silence, monitor, ask, or intervene. When action is warranted, they also provide a concrete response and a brief rationale.

silencemonitoraskintervene
  1. 01Engagement timing
  2. 02Action prediction
  3. 03Response + rationale
POSITIVE CASES

Safety alarms · Door cues · Appliance states · Impacts · Human non-speech cues

HARD NEGATIVES

Weak evidence · Premature moments · Resolved events · No-action scenes

03 / THE DATA

One signal format.
Two distinct tasks.

All inputs are 10–90 second audio segments sampled at 16 kHz, with reproducible tooling from training to hidden evaluation.

PartitionTrack ITrack II
TrainingLabeled350 h300 h
DevelopmentLabeled15 h15 h
LeaderboardInput-only20 h15 h
70% HIDDEN MAIN30% GENERALIZATION

Hidden Main follows the target-condition distribution while remaining disjoint in speakers, sessions, and source recordings.

Hidden Generalization introduces predefined shifts in speakers, environments, devices, sound sources, interaction patterns, and acoustic conditions.

Private evaluation · Submitted OCI containers · Frozen scorers · No network access

04 / EVALUATION

Independent rankings.
Transparent weights.

Each component is normalized to [0, 1]. A single score is computed for each track over the combined hidden benchmark.

TRACK I SCORESI

0.30 Decision + 0.35 Timing + 0.35 Action

30

Turn-taking decisionF1 for whether to take the turn.

35

Response timingAccuracy within temporal tolerance.

35

Action selectionTop-1 accuracy for the correct action.

TRACK II SCORESII

0.30 Timing + 0.30 Action + 0.40 Response

30

Engagement timingAccuracy of an appropriate engagement time.

30

Action predictionTop-1 accuracy over four actions.

40

Response + rationaleCorrectness and appropriateness.

TIE-BREAK Hidden Generalization performance, then lower response latency for Track I or lower false-activation rate for Track II.

05 / PARTICIPATION

Compete openly.
Evaluate reproducibly.

Public data and pretrained models are welcome when declared and legally usable. Award-eligible systems are evaluated under a transparent, frozen protocol.

ELIGIBILITY

Organizers and their PhD students are ineligible. Other institutional conflicts must be disclosed.

EXTERNAL RESOURCES

Public data and pretrained models must be declared, legally usable, and obtainable by the resource freeze.

SUBMISSIONS

Two development submissions per team/day; one primary and one backup final container.

REPRODUCIBILITY

Finalists provide a method card, manifests, model hashes, seeds, dependencies, and rerunnable code.

EXECUTION

No-network OCI containers run on randomized filenames with state reset between episodes.

NOT AWARD-ELIGIBLE

Private data, private test annotation, and private or commercial inference services.

06 / SCHEDULE

From launch to
ICASSP-ready.

All deadlines are at 23:59 UTC unless stated otherwise. The final container deadline is highlighted.

  1. SEP 072026

    Registration and website open

  2. SEP 212026

    Data, rules, scorers, and baselines released

  3. SEP 282026

    Development leaderboard opens

  4. NOV 202026

    Teams and resource declarations freeze

  5. DEC 042026

    Development leaderboard freezes

  6. DEC 112026

    Final container submission deadline

  7. DEC 182026

    Rankings and paper invitations announced

  8. JAN 072027

    Invited 2-page papers due

  9. JAN 282027

    Camera-ready 2-page papers due

07 / RECOGNITION

Two track winners.
Up to five invited teams.

05

Up to five audited, top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person at the dedicated ICASSP 2027 SPGC session.

Invitations consider final rank, technical diversity, reproducibility, and balanced coverage of the two tracks.
08 / ORGANIZERS

Built across research,
evaluation, and operations.

A cross-institutional team spanning audio research, benchmark design, data governance, infrastructure, and participant support.

GENERAL CHAIR

Renhe Sun

Conference liaison, governance, conflict management, and paper and session coordination.

DATASET & EVALUATION LEAD

Zihang Liu

Dataset construction, annotation, privacy, task design, metrics, baselines, and hidden-test auditing.

PLATFORM & OUTREACH LEAD

Jincenzi Wu

Website, evaluation platform, containers, compute, participant support, webinars, and outreach.

01Renhe Sun02Zihang Liu03Jincenzi Wu04Malu Zhang05Yuge Huang06Xiangdong Wang07Wenxuan Zhang08Jian Liu09Shuicheng Yan

Ant Group · National University of Singapore · The Chinese University of Hong Kong · University of Electronic Science and Technology of China · Institute of Computing Technology, Chinese Academy of Sciences · Singapore University of Technology and Design

09 / FAQ

Good questions.
Clear rules.

Are the two tracks combined?+

No. They share infrastructure and a schedule but have independent tasks, outputs, rankings, and winners.

What audio does HEARTBEAT use?+

Both tracks use 10–90 second segments sampled at 16 kHz. Track I focuses on conversational speech; Track II focuses on environmental and human non-speech audio.

Can we use public datasets or pretrained models?+

Yes, if they are declared, legally usable, and obtainable by the November 20 resource-freeze date.

How are final systems evaluated?+

Organizers execute submitted no-network OCI containers against private Hidden Main and Hidden Generalization benchmarks using frozen scorers.

What happens after the final ranking?+

Each track has a winner, and up to five audited top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person.

HEARTBEAT CHALLENGE 2027

Listen. Decide. Respond—or don't.

Registration, documentation, leaderboard access, tutorials, webinars, and the participant forum launch September 7, 2026.

Save the dates