ICASSP 2027 · SIGNAL PROCESSING GRAND CHALLENGE

Teach interactive agents when to respond— and when not to.

HEARTBEAT evaluates whether human-centered systems can decide whether, when, and how to respond to conversational audio and continuous audio-visual or audio-only streams.

STATUS / REGISTRATION OPEN Registration has been open since October 7, 2026

LIVE STREAM CONTEXT00:42.8
DECISION WINDOWContinue listening
92.4%
RECEIVEUPDATE STATEDECIDERESPOND

Only information received up to the current time may be used.

02Independent tracks
02Track I settings
02Track II subtracks
03Cognitive levels
01 / THE QUESTION

Listening is not enough. An agent must know what to do next.

Human-centered systems must distinguish genuine opportunities to respond from pauses, backchannels, self-repairs, ambiguous events, and moments where continued listening is better.

Knowing when not to act is part of acting intelligently.

01

Timely

Respond at the right moment—not merely as soon as possible.

02

Context-aware

Choose behavior that fits the conversation, current stream, and user goal.

03

Restrained

Remain silent when the current evidence does not require a response.

02 / THE TRACKS

Two tracks.
Multiple settings.

Track I has separate Offline and Online evaluations and rankings. Track II offers Audio-Visual and Audio-Only subtracks, both evaluated through online streaming.

TRACK I01

Reciprocal Audio Turn-Taking

Know when to begin responding—and when to stop speaking.

Systems participate naturally in ongoing conversations by using the audio context and their current state to decide whether to respond, keep listening, stop speaking, or handle new information appropriately.

Offline decision-makingOnline duplex interaction
  1. 01Offline · Four-way action selection
  2. 02Online · Response onset
  3. 03Online · Speech cessation
INTERACTION COVERAGE

Requests · Clarification · Correction · Information updates · Interruptions · Pause, cancel, and resume

CHALLENGING MOMENTS

Natural pauses · Hesitations · Self-repairs · Incomplete utterances · Backchannels · Side speech

TRACK II02

Streaming Proactive Understanding

Decide when to respond and what to say as the stream unfolds.

Systems receive one instruction, then process audio-visual or audio-only input online and in chronological order. They must use only information observed so far, support multiple responses in one stream, and remain silent when no response is needed.

Audio-visual subtrackAudio-only subtrackOnline only
  1. 01Perception
  2. 02Comprehension
  3. 03Reasoning
CAUSAL PROCESSING

Current and previous chunks only · Continuous state · No access to the complete recording

RESPONSE POLICY

Autonomous trigger detection · Multiple responses · Explicit silence when no response is needed

03 / THE PROTOCOL

Causal input.
Observable output.

Online systems process each stream in chronological order and may access only the current chunk and information received earlier.

SettingInputRequired output
Track I · OfflineOne prediction per sample

InputConversational audio prefix, assistant instructions, and four candidate actions

Required outputOne selected option per sample in JSON

Track I · OnlineReal-time duplex interaction

InputReal-time chunks of user speech, other participants' speech, and environmental audio

Required outputOutput audio plus JSON logs with chunk intervals, speech onset and offset, and response content

Track II · OnlineParticipant-selected chunks

InputAudio-visual or audio-only chunks received in chronological order

Required outputA JSON record for every chunk with its time interval and model output

CAUSAL ACCESS ONLYMULTIPLE RESPONSESEXPLICIT SILENCE

Continuous state Systems maintain context as new input arrives. They cannot inspect the full recording in advance or defer every response until the stream ends.

Time-accounted output For Track II, a response is timestamped at the end of the chunk that produced it, so chunk length directly affects measured latency.

Every Track II chunk must be logged, including chunks with no response.

04 / EVALUATION

Mode-specific scoring.
Timing that counts.

Each track follows its own task protocol. Correct content matters only when the system also responds at an appropriate time.

TRACK I · OFFLINEI / OFF

TOP-1 ACCURACY

01

Four-way selectionSelect one candidate action from the audio context at the designated decision point.

02

Exact scoringIncorrect, missing, or unparseable predictions are counted as errors.

03

No latency termOffline evaluation measures handling correctness only.

TRACK I · ONLINEI / ON

BEHAVIOR + QUALITY + TIMING

01

Turn-taking decisionCorrectly begin responding or stop speaking; missed, false, and duplicate triggers affect the result.

02

Response qualitySatisfy the request and context; different valid formulations are accepted.

03

Response timingMatch actual speech onset and cessation to their timing windows and report both latencies.

TRACK II · ONLINEII / ON

PRECISION / RECALL / JOINT F1

01

Temporal matchingMatch actual responses to ground-truth triggers inside the specified tolerance window.

02

Content correctnessVerify structured answers by task rules and open-ended responses against references.

03

Error sensitivityMissed events, false activations, and duplicate responses all reduce the score.

TRACK II TIMING A response is timestamped at the end of the chunk that generated it and cannot be backdated; chunk length therefore directly affects measured latency.

05 / PARTICIPATION

Compete openly.
Evaluate reproducibly.

Public data and pretrained models are welcome when declared and legally usable. Award-eligible systems are evaluated under a transparent, frozen protocol.

ELIGIBILITY

Organizers and their PhD students are ineligible. Other institutional conflicts must be disclosed.

EXTERNAL RESOURCES

Public data and pretrained models must be declared, legally usable, and obtainable by the resource freeze.

SUBMISSIONS

Two development submissions per team/day; one primary and one backup final container.

REPRODUCIBILITY

Finalists provide a method card, manifests, model hashes, seeds, dependencies, and rerunnable code.

EXECUTION

No-network OCI containers run on randomized filenames with state reset between episodes.

NOT AWARD-ELIGIBLE

Private data, private test annotation, and private or commercial inference services.

06 / SCHEDULE

From launch to
ICASSP-ready.

All deadlines: 23:59 UTC unless otherwise stated. The final container deadline is highlighted.

  1. OCT 072026

    Registration opens and website goes live

  2. OCT 142026

    Data, rules, scoring tools, and baseline systems released

  3. OCT 212026

    Development leaderboard opens

  4. NOV 202026

    Team membership and resource declarations frozen

  5. DEC 042026

    Development leaderboard closes

  6. DEC 112026

    Final container submission deadline

  7. DEC 182026

    Final rankings and paper invitations announced

  8. JAN 072027

    Invited two-page papers due

  9. JAN 282027

    Camera-ready two-page papers due

07 / RECOGNITION

Two track winners.
Up to five invited teams.

05

Up to five audited, top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person at the dedicated ICASSP 2027 SPGC session.

Invitations consider final rank, technical diversity, reproducibility, and balanced coverage of the two tracks.
08 / ORGANIZERS

Built across research,
evaluation, and operations.

A cross-institutional team spanning audio research, benchmark design, data governance, infrastructure, and participant support.

GENERAL CHAIR

Renhe Sun

Conference liaison, governance, conflict management, and paper and session coordination.

DATASET & EVALUATION LEAD

Zihang Liu

Dataset construction, annotation, privacy, task design, metrics, baselines, and hidden-test auditing.

PLATFORM & OUTREACH LEAD

Jincenzi Wu

Website, evaluation platform, containers, compute, participant support, webinars, and outreach.

01Renhe Sun02Zihang Liu03Jincenzi Wu04Malu Zhang05Yuge Huang06Xiangdong Wang07Wenxuan Zhang08Jian Liu09Shuicheng Yan

Ant Group · National University of Singapore · The Chinese University of Hong Kong · University of Electronic Science and Technology of China · Institute of Computing Technology, Chinese Academy of Sciences · Singapore University of Technology and Design

09 / FAQ

Good questions.
Clear rules.

Are the two tracks combined?+

No. Track I has separate Offline and Online evaluations and rankings. Track II has Audio-Visual and Audio-Only subtracks, and participants may register for either subtrack.

What input does each track use?+

Track I uses conversational audio. Track II offers an audio-visual subtrack and an audio-only subtrack; participants may enter either one.

What does online processing allow?+

Systems process inputs in chronological order and may use only the current and previously received information. They cannot inspect the complete recording in advance or wait until the end to produce every answer.

Can we use public datasets or pretrained models?+

Yes, if they are declared, legally usable, and obtainable by the November 20 resource-freeze date.

What do participants submit?+

Track I Offline requires a JSON file with sample identifiers and selected options. Track I Online requires output audio and timestamped JSON logs. Track II requires per-chunk JSON containing time ranges and model outputs, including chunks with no response.

How are systems evaluated?+

Track I Offline uses Top-1 accuracy; Track I Online evaluates turn-taking behavior, response quality, and response timing. Track II uses Precision, Recall, and joint F1 over correctly timed and correct responses.

What happens after the final ranking?+

Each track has a winner, and up to five audited top-ranked teams across both rankings will be invited to submit 2-page ICASSP papers and present in person.

HEARTBEAT CHALLENGE 2027

Listen. Decide. Respond—or don't.

Registration is open. Complete the official form to enter HEARTBEAT Challenge 2027. Data, rules, scoring tools, and baseline systems follow October 14; the development leaderboard opens October 21.

Register now