Does Consciousness Arrive in Frames?
창밖을 지나가는 자동차를 바라보면 움직임은 끊기지 않는다. 엔진 소리는 멀리서 다가왔다가 가까워지고 다시 희미해진다. 시각과 청각은 하나의 연속된 장면을 만들어 낸다. 경험만 놓고 보면 세계는 흐른다.
그러나 실험실에서 지각을 측정하면 이 매끄러운 인상과 잘 맞지 않는 현상들이 나타난다. 같은 밝기의 희미한 섬광도 어떤 순간에는 보이고 불과 수십 밀리초 뒤에는 보이지 않는다. 자극에 대한 초기 신경 반응은 점진적으로 변하지만, 참가자의 보고는 어느 지점에서 갑자기 보았다로 전환된다. 때로는 뒤늦게 들어온 정보가 조금 전에 본 것의 모습까지 바꾼다. 이런 결과는 오래된 질문을 되살린다. 의식은 정말 연속적인가, 아니면 영화처럼 일련의 프레임으로 구성되는가.
지금까지의 연구는 이 질문에 단순한 양자택일로 답하지 않는다. 뇌에는 모든 경험을 같은 속도로 잘라 내는 하나의 셔터가 있는 것 같지 않다. 대신 서로 다른 시간 구조가 겹쳐 있다. 신경 활동은 계속 변하고, 감각 선택은 뇌의 리듬에 따라 흔들리며, 정보가 보고와 행동에 사용 가능해지는 과정에는 문턱이 있다. 지각 내용은 짧은 시간 동안 통합되고, 중요한 변화가 생길 때 새로운 상태로 재편된다. 그러니 핵심은 의식이 연속인가 이산인가를 고르는 데 있지 않다. 어느 수준이 연속적이고 어느 수준에서 전환이 일어나는지를 구분하는 데 있다.
여기서 다루는 의식은 주로 시각적 내용이 보고와 작업기억과 계획과 행동에 사용 가능해지는 과정, 곧 접근 의식에 가깝다. 경험 그 자체인 현상적 의식은 직접 측정하기 어렵다. 실험이 기록하는 것은 대개 버튼 누르기, 언어 보고, 기억 성적, 뇌 신호 같은 간접 지표다. 이 제한을 무시하면 행동의 이산성을 경험 자체의 이산성으로 오해하기 쉽다.[1]
의식의 이산성을 논할 때는 서로 다른 네 가지 문제를 분리해야 한다. 하나는 신경 과정의 시간적 형태에 관한 것이다. 뉴런 집단의 막전위와 발화율, 영역 사이의 상호작용이 매 순간 연속적으로 변하는지, 아니면 특정 시점에만 갱신되는지를 묻는다. 다른 하나는 감각 선택의 주기성이다. 뇌가 모든 입력을 같은 강도로 처리하지 않고 일정한 리듬에 따라 어떤 정보는 증폭하고 다른 정보는 억제하는지를 묻는다. 세 번째는 접근의 문턱이다. 자극에 관한 증거가 점진적으로 쌓이다가 특정 수준을 넘는 순간 그 정보가 보고와 작업기억에 갑자기 들어오는지를 묻는다. 마지막은 내용의 사건성이다. 현재의 경험이 잠시 안정적으로 유지되다가 객체가 바뀌거나 목표가 달라지거나 예상이 빗나가는 순간 새로운 상태로 재구성되는지를 묻는다.
이 네 문제는 서로 연관되어 있지만 동일하지 않다. 행동 보고가 보았다와 보지 못했다로 나뉜다고 해서 그 이전의 신경 과정까지 이진적이었다고 말할 수는 없다. 탐지 성능이 주기적으로 흔들린다고 해서 주기 사이에 경험이 사라진다고 결론 내릴 수도 없다. 사건 경계에서 기억이 분절된다고 해서 뇌가 일정한 간격의 프레임을 사용한다는 뜻도 아니다. 지각 성능이 리듬에 따라 변한다는 것은 실험적 관찰이고, 의식이 리듬의 골짜기에서 꺼진다는 것은 그보다 훨씬 강한 해석이다. 전자는 여러 연구가 지지하지만 후자는 아직 입증되지 않았다. 이 구분을 염두에 두면 의식의 시간 구조에 관한 증거들이 서로 모순되지 않고 한 그림 안에 들어오기 시작한다.
뇌는 리듬을 가진 기관이다. 알파파와 세타파를 비롯한 여러 진동은 감각피질과 두정엽과 전두엽 사이의 정보 교환과 관련된다. 이 리듬은 단순한 배경 소음이 아니다. 자극이 나타나기 직전의 진동 위상은 역치 부근 시각 자극이 보일 가능성을 예측하며, 여러 위치에 주의를 나눌 때에는 탐지 성능이 수 헤르츠 범위에서 번갈아 높아지는 현상이 관찰된다.[2-4] Landau와 Fries의 실험에서 참가자가 두 위치 중 어디에 표적이 나타날지 감시할 때, 두 위치의 탐지 성능은 같은 수준으로 유지되지 않고 한쪽의 민감도가 올라갈 때 다른 쪽이 내려가는 식으로 주의의 우선순위가 교대했다.[3] Fiebelkorn과 동료들의 원숭이 연구는 이런 리듬이 감각 영역 하나의 고립된 진동이라기보다 전두엽과 두정엽 네트워크의 동적 상호작용과 관련될 수 있음을 보였다.[5]
이 결과만 보면 뇌가 세상을 연속적으로 바라보는 대신 주기적으로 표본화한다고 생각하기 쉽다. 그러나 여기서 측정된 것은 경험의 존재 여부가 아니라 처리 효율이다. 같은 자극이라도 신경계의 흥분성과 주의 이득이 높은 위상에서는 더 잘 탐지되고 낮은 위상에서는 놓칠 가능성이 커진다. 카메라의 셔터는 노출 사이에 빛을 완전히 차단하지만, 뇌의 리듬은 조명 조절기에 가깝다. 입력은 계속 들어오되 어떤 순간에는 특정 정보의 영향력이 커지고 다른 순간에는 작아진다. 감각 증거 자체는 연속적으로 변하고, 그 증거가 접근에 기여하는 가중치가 주기적으로 달라지는 것이다. 게다가 지각과 주의에서 관찰되는 주파수는 하나가 아니다. 시각적 시간 해상도는 개인의 알파 주파수와 관련되기도 하고, 여러 대상을 오가는 주의는 더 느린 세타 범위에서 나타나기도 한다.[2,4] 만약 의식 전체를 지배하는 단일한 프레임레이트가 있다면 서로 다른 과제에서 공통된 기본 주기가 반복적으로 드러나야 하는데, 현재까지의 자료는 그런 보편적 시계보다 여러 회로가 서로 다른 시간척도로 작동한다는 해석에 더 잘 맞는다. 행동 자료에서 주기성을 찾는 분석은 통계적으로도 조심해야 한다. 자극이 한 번 일으킨 과도 반응이나 시간적 기대가 반복 진동처럼 보일 수 있기 때문이다. 리듬은 의식의 시간을 조직하는 한 요소일 수 있다. 그러나 그것이 곧 의식의 프레임은 아니다.
희미한 자극의 밝기를 조금씩 높여 가면 주관적 보고는 흔히 완만하게 증가하지 않는다. 어느 구간까지는 거의 보이지 않다가 좁은 범위에서 보임의 비율이 빠르게 올라간다. 이 급격한 전환은 의식적 접근에 문턱이 있을 가능성을 제시한다. 역행 마스킹은 이 현상을 연구하는 대표적 방법이다. 표적을 매우 짧게 보여 준 뒤 곧바로 다른 자극으로 덮으면 초기 시각계는 표적에 반응하면서도 참가자는 그것을 보고하지 못할 수 있다. Del Cul과 Baillet과 Dehaene의 MEG 연구에서 보고되지 않은 표적도 초기 후두-측두 경로에서 처리되었지만, 의식적으로 보고된 시행에서는 약 270밀리초 이후 더 넓은 피질 네트워크가 관여하는 후기 활동이 나타났다.[6]
이 결과는 연속적인 과정이 갑작스러운 출력을 만드는 방식을 보여 준다. 감각 증거는 시간에 따라 축적되고 재귀적 상호작용을 통해 증폭될 수 있다. 증거가 충분하지 않으면 초기 처리에 머물지만, 일정 수준을 넘으면 작업기억과 언어 보고와 계획에 널리 사용 가능한 상태로 전환된다. 물이 서서히 차오르다가 둑을 넘는 순간을 생각하면 된다. 수위의 변화는 연속적이지만 넘침은 사건처럼 보인다. 다만 여기에는 주의점이 있다. 후기의 광범위한 활동이 의식 자체를 뜻한다고 단정할 수는 없다. 참가자가 무엇을 보았다고 보고하려면 내용을 유지하고 결정을 내리고 운동 반응을 준비해야 하는데, 후기 신호에는 이런 과제 요구가 섞인다. 광범위한 점화가 경험의 원인인지 보고의 결과인지 둘을 함께 반영하는지 분리해야 한다. 2025년 Cogitate Consortium의 적대적 협업 연구가 이 문제를 정면으로 다루었다. 연구진은 통합정보이론과 전역 신경 작업공간 이론이 서로 다르게 예측하는 결과를 사전에 정하고 256명에게 fMRI와 MEG와 두개내 EEG를 적용했다. 의식 내용과 관련된 정보는 후두·복측두 영역뿐 아니라 일부 전두 영역에서도 관찰되었고 후방 피질의 일부 반응은 자극이 제시되는 동안 지속되었지만, 두 이론의 강한 예측은 모두 일부 반박되었다.[7] 이 결과는 의식이 오직 한 번의 전역적 폭발로 생긴다고 보기 어렵게 만든다. 접근에는 비선형적 전환이 있을 수 있지만 내용 자체는 여러 영역에서 지속적으로 유지될 수 있다. 전환과 지속은 서로 배타적이지 않다.
우리는 지금을 순간처럼 느끼지만 지각이 사용하는 현재는 수학적인 한 점이 아니다. 뇌는 짧은 시간 동안 들어온 정보를 묶어 하나의 사건으로 만든다. 그 사실은 후시적 구성에서 뚜렷하게 드러난다. 어떤 자극을 본 뒤 수십에서 수백 밀리초 안에 들어온 정보가 앞선 자극의 최종 지각을 바꿀 수 있다. 나중의 사건이 이전의 신경 반응을 시간 여행하듯 바꾸는 것은 아니다. 오히려 앞선 자극의 지각이 처음부터 완전히 확정되지 않았다고 보는 편이 자연스럽다. 뇌는 잠시 여러 해석을 유지한 채 뒤이어 들어오는 단서를 사용해 하나의 결과를 구성한다. Herzog와 Kammer와 Scharnowski는 이 현상을 설명하려고 시간 조각 모델을 제안했다. 특징 분석은 높은 시간 해상도로 준연속적으로 진행되지만 의식적 지각은 일정한 통합 구간을 거쳐 형성된다는 생각이다.[8] 이 모델에서 우리가 경험하는 것은 매 순간의 원시 신호가 아니라 짧은 구간에 걸친 처리 결과를 압축한 내용이다.
음악을 들을 때 하나의 음은 앞뒤 음과 분리되어 의미를 갖지 않는다. 운동 방향도 최소한 두 시점의 위치 변화가 있어야 정의되고, 언어의 한 단어 역시 앞선 문맥과 뒤따르는 단어에 따라 해석이 달라진다. 시간 통합은 의식의 예외적 기능이 아니라 지각이 의미를 구성하기 위한 기본 조건이다. 그렇다고 통합창이 곧 비중첩 프레임이라는 뜻은 아니다. 시간창은 서로 겹칠 수 있고 정보의 종류에 따라 길이가 다를 수 있다. 움직임과 음성과 문장과 사회적 사건은 서로 다른 시간 규모를 요구한다. 연속적인 재귀 동역학이나 점진적인 확률 갱신도 후시적 효과를 설명할 수 있다. 실제로 이산 지각 이론에 대한 비판은 현재의 실험이 시간적 통합을 지지할 뿐 프레임 사이의 무경험 구간까지 입증하지는 않는다고 지적한다.[9] 이 단계에서 말할 수 있는 것은 제한적이지만 분명하다. 지각은 입력 순간에 즉시 완성되지 않는다. 현재의 내용은 짧은 과거를 포함하며, 그 길이는 하나의 보편적 숫자로 고정되어 있지 않다.
연속적인 장면을 보더라도 기억은 모든 순간을 같은 밀도로 저장하지 않는다. 사람이 방에 들어오고 컵을 집고 물을 따른 뒤 자리를 떠나는 장면을 보면 우리는 그것을 무한히 많은 자세의 연속으로 기억하지 않는다. 들어옴, 컵을 집음, 물을 따름, 나감 같은 사건으로 나눈다. 사건 분절 이론은 뇌가 현재 상황에 대한 내부 모형을 유지하다가 예측이 크게 빗나가는 순간 그 모형을 갱신한다고 본다.[10] 행동 목표가 바뀌거나 공간이 달라지거나 새로운 인물이 등장하거나 인과 구조가 전환되면 사건 경계가 형성되고, 이 경계는 이후 기억의 단위에도 영향을 준다.
사건 경계는 의식의 이산성을 논할 때 중요한 단서를 준다. 의식 내용의 큰 변화가 일정한 간격으로 일어나는 것이 아니라 정보 구조가 달라지는 순간에 일어난다면, 의식의 시간은 시계보다 사건에 의해 조직된다고 볼 수 있다. 이때의 경계는 고정되어 있지 않다. 같은 영상을 보더라도 요리법을 배우는 사람과 등장인물의 감정을 추적하는 사람은 서로 다른 순간을 중요하게 여긴다. 사건은 여러 규모로 중첩된다. 손을 뻗는 동작은 짧은 사건이고, 한 끼를 준비하는 과정은 더 긴 사건이며, 식사를 대접하는 사회적 맥락은 그보다 더 긴 사건이다. 이런 계층성과 과제 의존성은 고정 프레임 가설과 잘 맞지 않는다. 대신 연속적인 신경 과정 위에 여러 규모의 준안정 상태가 형성되고, 예측오차와 목표 관련성이 커질 때 상태가 전환된다는 그림을 지지한다.
의식의 시간 구조를 설명하는 계산 모델은 크게 셋으로 나눌 수 있다. 연속 모델에서는 내용이 입력과 함께 매 순간 조금씩 변한다. 고정 프레임 모델에서는 일정한 주기마다 전체 내용이 갱신된다. 하이브리드 모델에서는 기저 처리가 연속적으로 진행되지만 사건이 발생할 때 안정된 내용이 새로운 상태로 전환된다. 현재의 증거를 가장 적은 가정으로 묶으면 세 번째 모델이 유력하다. 감각 입력과 신경 상태는 계속 변하고, 주의 리듬은 특정 정보의 이득을 조절하며, 최근 정보는 짧은 시간창 안에서 통합되고, 예측오차와 현저성과 목표 관련성이 충분히 커지면 현재 내용이 새로운 준안정 상태로 재구성된다. 계산적으로는 이렇게 요약할 수 있다.
dz/dt = f(z(t), x(t))
c(k + 1) = B(z[t_k - tau : t_k])
when g(z(t), x(t), m(t)) > theta여기서 z(t)는 연속적으로 변하는 신경 상태이고 x(t)는 감각 입력이며 m(t)는 주의와 목표와 예측오차 같은 조절 변수다. 함수 g가 문턱 theta를 넘으면 최근 시간창의 정보가 새로운 내용 c(k + 1)로 조직된다. 중요한 것은 갱신 시점 t_k가 일정하지 않다는 점이다. 고정 프레임 모델은 t_k = kT인 특수한 경우다. 반면 사건 기반 모델에서는 조용하고 예측 가능한 장면에서 한 상태가 오래 유지될 수 있고, 변화가 많은 장면에서는 짧은 간격으로 여러 번 갱신될 수 있다. 이 모델은 의식이 프레임처럼 보이는 이유와 실제로 연속적으로 느껴지는 이유를 동시에 설명한다. 내용은 일정 시간 안정되어 있으므로 하나의 장면처럼 경험되지만, 그 아래에서는 감각 증거와 내부 상태가 계속 변한다. 이산적인 것은 신경 과정의 존재가 아니라 내용의 재조직화다. 하이브리드 모델이 옳다고 확정할 수는 없다. 그러나 알려진 리듬적 선택과 비선형적 접근과 시간 통합과 사건 분절을 한 구조 안에 넣을 수 있다는 점에서 가장 경제적인 작업 가설이다.
인공지능 모델을 의식 연구에 끌어올 때 가장 먼저 할 일은 경계를 긋는 것이다. 모델이 정보를 압축하고 기억하고 사건을 구분한다고 해서 주관적 경험을 가진다는 뜻은 아니다. 현재 AI 시스템의 현상적 의식을 판정하는 데 합의된 실험은 없으며, 계산 기능의 유사성은 경험의 동일성을 보장하지 않는다.[11] 그럼에도 AI는 의식 이론의 계산적 구성요소를 분리해 시험하는 데 유용하다. 인간과 동물의 뇌에서는 선택과 통합과 기억과 행동이 한꺼번에 일어나지만, 인공 시스템에서는 각 기능을 따로 구현하고 제거할 수 있다. 어떤 병목이 필요한지, 갱신 주기를 고정해야 하는지, 사건 경계에서만 기억하는 것이 유리한지를 직접 비교할 수 있다.
DeepMind의 Perceiver는 거대한 감각 입력을 작은 잠재 배열로 반복 압축한다.[12] 모든 입력이 같은 비중으로 깊은 계산에 참여하는 것이 아니라 제한된 잠재 공간이 선택적으로 정보를 받아들인다. 병목이 있다고 해서 입력 처리가 프레임 단위일 필요는 없다는 것을 보여주는 구조다. Allen Institute의 MERLOT Reserve는 비디오와 언어와 소리를 함께 학습해 시간적으로 연결된 사건 구조를 추론한다.[13] 픽셀과 음향의 연속을 그대로 저장하기보다 무엇이 먼저 일어났고 무엇이 다음에 일어날지를 예측하는 데 유용한 표현을 형성한다. 연속 스트림이 사건 단위의 내부 표상으로 압축될 수 있다는 계산적 사례다. DeepMind의 Differentiable Neural Computer와 Neural Episodic Control은 또 다른 시간 분리를 보여 준다.[14,15] 하나는 외부 메모리에 선택적으로 읽고 쓰며, 다른 하나는 느리게 변하는 신경망 가중치와 빠르게 갱신되는 에피소드 기억을 분리한다. 감각 처리와 현재 상태와 사건 기억과 장기 학습이 서로 다른 속도로 움직일 수 있다는 뜻이다. ALFRED 같은 체화 에이전트 벤치마크에서는 과거 행동이 환경을 바꾸고 그 변화가 다음 행동의 조건이 된다.[16] 에이전트는 현재 화면만 분류해서는 작업을 끝낼 수 없고, 무엇을 이미 집었는지 어떤 문이 열렸는지 어떤 목표가 완료되었는지를 사건 단위로 추적해야 한다.
이 모델들이 의식을 설명하는 것은 아니다. 그러나 의식 이론이 암묵적으로 요구하는 계산을 명시적인 부품으로 바꾸어 준다. 가장 직접적인 실험은 동일한 데이터와 계산 예산으로 세 에이전트를 비교하는 것이다. 하나는 매 순간 상태를 갱신하고, 하나는 일정한 주기로만 갱신하며, 하나는 예측오차나 목표 관련성이 문턱을 넘을 때 갱신한다. 장기 예측과 사건 기억과 행동 성공률과 계산 비용을 함께 측정하면 어떤 시간 구조가 적응적으로 유리한지 알 수 있다. 이런 실험은 AI가 의식적인지를 답하지 않는다. 대신 의식과 비슷한 기능을 만들려면 어떤 시간 구조가 필요한가라는 더 제한적이고 검증 가능한 질문에 답한다.
연속 모델과 고정 프레임 모델과 사건 기반 하이브리드 모델은 현재의 많은 결과를 사후적으로 설명할 수 있다. 이들을 구분하려면 각 모델이 실패할 수 있는 실험이 필요하다. 초기 감각 표상은 자극 강도에 따라 매끄럽게 변하지만 작업기억과 보고 가능성은 특정 시점에 급격히 전환되는지를 같은 시행에서 분리해 측정해야 하고, 같은 시간이 흘러도 예측 가능한 장면에서는 상태 전환이 적고 인과 구조가 바뀌는 장면에서는 전환이 많아지는지를 확인해 갱신이 시계에 정렬되는지 예측오차에 정렬되는지 보아야 한다. 의식 내용과 관련된 신호가 보고 준비와 운동 계획 때문에 나타나는 것은 아닌지 보고가 필요한 조건과 필요하지 않은 조건을 비교해야 하고, 행동 성능의 주기성만 찾는 대신 특정 리듬의 위상을 자극 시점과 정렬하거나 비침습적 자극으로 리듬을 교란했을 때 지각의 시간 구조가 예측대로 바뀌는지 살펴야 한다. 무엇보다 세 모델을 같은 자료에 적합한 뒤 보지 않은 시행에서 상태 전환 시점과 보고를 얼마나 정확히 예측하는지를 비교해야 한다. 의식 연구에서 그럴듯한 서사는 많다. 모델을 가르는 것은 새로운 데이터에 대한 예측이다.
현재의 증거는 의식이 일정한 속도로 재생되는 영화라는 생각을 지지하지 않는다. 모든 감각과 사고를 하나의 프레임레이트로 묶는 보편적 시계도 발견되지 않았다. 반대로 의식의 모든 측면이 매끄럽게 변한다고 말하기도 어렵다. 지각 민감도는 리듬을 타고, 접근은 문턱을 가지며, 내용은 시간에 걸쳐 구성되고, 사건 경계에서 빠르게 재편된다. 그러니 의식은 이산적인가라는 질문에는 수준별로 답해야 한다. 신경 동역학은 대체로 연속적이다. 정보 선택은 주기적으로 변조될 수 있다. 보고와 작업기억으로의 접근은 비선형적일 수 있다. 의식 내용은 준안정적으로 유지되다가 사건에 따라 갱신될 수 있다. 그러나 경험 자체가 그 사이에 완전히 사라지는지는 아직 알 수 없다.
가장 보수적인 결론은 이렇다. 의식은 고정된 프레임의 연속이라기보다, 흐르는 신경 과정이 짧은 시간 동안 하나의 내용으로 조직되고 중요한 사건 앞에서 새로운 상태로 넘어가는 체계에 가깝다. 이 관점에서 순간은 의식의 최소 입자가 아니다. 연속적인 과정이 잠시 안정된 형태를 얻는 방식이다.
참고문헌
- Block, N. (1995). On a Confusion about a Function of Consciousness. Behavioral and Brain Sciences, 18(2), 227-247. doi:10.1017/S0140525X00038188
- VanRullen, R. (2016). Perceptual Cycles. Trends in Cognitive Sciences, 20(10), 723-735. doi:10.1016/j.tics.2016.07.006
- Landau, A. N., & Fries, P. (2012). Attention Samples Stimuli Rhythmically. Current Biology, 22(11), 1000-1004. doi:10.1016/j.cub.2012.03.054
- Busch, N. A., Dubois, J., & VanRullen, R. (2009). The Phase of Ongoing EEG Oscillations Predicts Visual Perception. Journal of Neuroscience, 29(24), 7869-7876. doi:10.1523/JNEUROSCI.0113-09.2009
- Fiebelkorn, I. C. et al. (2018). A Dynamic Interplay within the Frontoparietal Network Underlies Rhythmic Spatial Attention. Neuron, 99(4), 842-853.e8. doi:10.1016/j.neuron.2018.07.038
- Del Cul, A., Baillet, S., & Dehaene, S. (2007). Brain Dynamics Underlying the Nonlinear Threshold for Access to Consciousness. PLOS Biology, 5(10), e260. doi:10.1371/journal.pbio.0050260
- Cogitate Consortium et al. (2025). Adversarial Testing of Global Neuronal Workspace and Integrated Information Theories of Consciousness. Nature, 642, 133-142. doi:10.1038/s41586-025-08888-1
- Herzog, M. H., Kammer, T., & Scharnowski, F. (2016). Time Slices: What Is the Duration of a Percept? PLOS Biology, 14(4), e1002433. doi:10.1371/journal.pbio.1002433
- Fekete, T., Van de Cruys, S., Ekroll, V., & van Leeuwen, C. (2018). In the Interest of Saving Time: A Critique of Discrete Perception. Neuroscience of Consciousness, 2018(1), niy003. doi:10.1093/nc/niy003
- Zacks, J. M. et al. (2007). Event Perception: A Mind-Brain Perspective. Psychological Bulletin, 133(2), 273-293. doi:10.1037/0033-2909.133.2.273
- Butlin, P. et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.
- Jaegle, A. et al. (2021). Perceiver: General Perception with Iterative Attention. Proceedings of ICML 2021, PMLR 139, 4651-4664.
- Zellers, R. et al. (2022). MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound. Proceedings of CVPR 2022.
- Graves, A. et al. (2016). Hybrid Computing Using a Neural Network with Dynamic External Memory. Nature, 538, 471-476. doi:10.1038/nature20101
- Pritzel, A. et al. (2017). Neural Episodic Control. Proceedings of ICML 2017, PMLR 70, 2827-2836.
- Shridhar, M. et al. (2020). ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. Proceedings of CVPR 2020, 10740-10749.
When you watch a car pass outside a window, its motion does not break apart. The engine sound approaches from a distance, comes closer, then fades again. Vision and hearing form a single continuous scene. As experience alone, the world flows.
But when perception is measured in the laboratory, phenomena appear that do not fit this smooth impression. A faint flash of the same brightness may be seen at one moment and missed only tens of milliseconds later. Early neural responses to a stimulus change gradually, but a participant's report switches abruptly at some point to "seen." Sometimes information that arrives late can even alter the appearance of what was seen a moment before. Results like these revive an old question. Is consciousness truly continuous, or is it composed of a sequence of frames, like film?
Research so far does not answer this question with a simple either-or. The brain does not seem to contain a single shutter that slices every experience at the same speed. Instead, different temporal structures overlap. Neural activity keeps changing, sensory selection fluctuates with the brain's rhythms, and the process by which information becomes available for report and action has thresholds. Perceptual content is integrated over short periods and reorganized into a new state when an important change occurs. The key, then, is not to choose between continuous and discrete. It is to distinguish which levels are continuous and where transitions occur.
The consciousness discussed here is mainly the process by which visual content becomes available for report, working memory, planning, and action — that is, access consciousness. Phenomenal consciousness, experience itself, is harder to measure directly. What experiments record are usually indirect indicators such as button presses, verbal reports, memory performance, or brain signals. If we ignore this limitation, it becomes easy to mistake discreteness in behavior for discreteness in experience itself.[1]
When discussing the discreteness of consciousness, four different questions have to be separated. One concerns the temporal form of neural processes: whether the membrane potentials and firing rates of neuronal populations, and the interactions between areas, change continuously at every moment or are updated only at specific times. Another is the periodicity of sensory selection — whether the brain processes all inputs with equal strength, or amplifies some information and suppresses other information according to a steady rhythm. A third is the threshold of access: whether evidence about a stimulus accumulates gradually and then, at the moment it crosses a certain level, enters report and working memory abruptly. The last is the event structure of content: whether current experience remains briefly stable and then reorganizes into a new state the moment an object changes, a goal shifts, or a prediction fails.
These four questions are related, but they are not the same. That behavioral reports divide into seen and not-seen does not let us say that the preceding neural process was also binary. That detection performance fluctuates periodically does not allow the conclusion that experience disappears between cycles. That memory is segmented at event boundaries does not mean the brain uses fixed-interval frames. That perceptual performance varies with rhythm is an experimental observation; that consciousness switches off in the troughs of the rhythm is a far stronger interpretation. The former is supported by multiple studies, but the latter has not yet been established. With this distinction in mind, the evidence about the temporal structure of consciousness begins to fit into one picture rather than contradicting itself.
The brain is a rhythmic organ. Oscillations such as alpha and theta are linked to the exchange of information between sensory cortex, parietal cortex, and frontal cortex. These rhythms are not mere background noise. The phase of an oscillation just before a stimulus appears predicts whether a near-threshold visual stimulus will be seen, and when attention is divided across multiple locations, detection performance has been observed to rise and fall alternately in the range of several hertz.[2-4] In the experiment by Landau and Fries, when participants monitored which of two locations would contain a target, detection performance at the two locations did not stay equal; as sensitivity rose at one location it fell at the other, so that attentional priority alternated.[3] Fiebelkorn and colleagues' work in monkeys showed that such a rhythm may involve dynamic interactions in the frontoparietal network rather than an isolated oscillation in a single sensory area.[5]
From these findings alone it is tempting to think that the brain samples the world periodically instead of viewing it continuously. But what was measured here is not the presence or absence of experience; it is processing efficiency. The same stimulus is better detected in phases where neural excitability and attentional gain are high, and more likely to be missed in phases where they are low. A camera shutter completely blocks light between exposures, but the brain's rhythm is closer to a dimmer. Input keeps arriving, yet at some moments the influence of particular information grows and at others it shrinks. Sensory evidence itself changes continuously, while the weight with which that evidence contributes to access changes periodically. Moreover, the frequencies observed in perception and attention are not singular. Visual temporal resolution can be related to an individual's alpha frequency, while attention moving among several objects can appear in the slower theta range.[2,4] If there were a single frame rate governing consciousness as a whole, a shared fundamental period should surface repeatedly across different tasks; the evidence so far fits better with the view that multiple circuits operate on different time scales than with such a universal clock. Analyses that search for periodicity in behavioral data also call for statistical caution, because a single transient response triggered by a stimulus, or a temporal expectation, can look like a repeated oscillation. Rhythm may be one element that organizes the time of consciousness. But it is not, by itself, the frame of consciousness.
When the brightness of a faint stimulus is raised little by little, subjective reports usually do not increase smoothly. For a range they are almost never seen, and then within a narrow band the proportion of "seen" reports rises quickly. This abrupt transition suggests that conscious access may have a threshold. Backward masking is a standard method for studying this phenomenon. If a target is shown very briefly and then immediately covered by another stimulus, the early visual system may respond to the target while the participant still cannot report it. In the MEG study by Del Cul, Baillet, and Dehaene, unreported targets were nonetheless processed in early occipito-temporal pathways, but in trials where the target was consciously reported, later activity involving a wider cortical network appeared after roughly 270 milliseconds.[6]
This result shows how a continuous process can produce a sudden output. Sensory evidence accumulates over time and can be amplified through recurrent interactions. If the evidence is insufficient, processing stays early; once it passes a certain level, it shifts into a state widely available to working memory, verbal report, and planning. Think of water rising slowly until it spills over a bank. The change in water level is continuous, but the overflow looks like an event. There is a caveat here, though. We cannot conclude that later widespread activity means consciousness itself. To report what they saw, participants must maintain the content, make a decision, and prepare a motor response, and these task demands are mixed into the late signal. We have to separate whether the widespread ignition is the cause of experience, the result of report, or a reflection of both. In 2025 the Cogitate Consortium's adversarial collaboration confronted this problem directly. The researchers preregistered the predictions on which Integrated Information Theory and Global Neuronal Workspace Theory differ, and applied fMRI, MEG, and intracranial EEG to 256 participants. Information related to conscious content was observed not only in occipital and ventral temporal regions but also in some frontal regions, and some posterior cortical responses persisted while the stimulus was present; yet the strong predictions of both theories were each partly contradicted.[7] This makes it hard to see consciousness as arising from a single global explosion. Access may involve a nonlinear transition, but content itself can be sustained across multiple areas. Transition and persistence are not mutually exclusive.
We feel "now" as an instant, but the present that perception uses is not a mathematical point. The brain binds information arriving over a short period into a single event. That fact shows clearly in postdictive construction. Information arriving within tens to hundreds of milliseconds after a stimulus can change the final perception of the earlier stimulus. It is not that the later event travels back in time and alters the earlier neural response. It is more natural to say that perception of the earlier stimulus was never fully settled from the start. The brain holds several interpretations for a moment and uses the cues that follow to construct one result. To explain this phenomenon, Herzog, Kammer, and Scharnowski proposed the time-slices model — the idea that feature analysis proceeds quasi-continuously at high temporal resolution, while conscious perception is formed through a fixed interval of integration.[8] In this model, what we experience is not the raw signal at every instant but content compressed from processing over a short interval.
When listening to music, a single note has no meaning apart from the notes before and after it. Motion direction, too, is defined only by a change in position across at least two moments, and a single word in language is interpreted differently depending on the preceding context and the words that follow. Temporal integration is not an exceptional function of consciousness but a basic condition for perception to construct meaning. Even so, an integration window is not the same as a set of non-overlapping frames. Temporal windows can overlap, and their length can differ by the type of information. Motion, speech, sentences, and social events demand different time scales. Continuous recurrent dynamics or gradual probabilistic updating can also account for postdictive effects. Indeed, critiques of discrete-perception theory point out that current experiments support temporal integration but do not prove gaps without experience between frames.[9] What can be said at this stage is limited but clear. Perception is not completed the moment input arrives. The content of the present includes a short past, and its length is not fixed by any single universal number.
Even when we watch a continuous scene, memory does not store every moment at the same density. Watching someone enter a room, pick up a cup, pour water, and leave, we do not remember it as an endless sequence of postures. We divide it into events: entering, picking up the cup, pouring water, leaving. Event Segmentation Theory holds that the brain maintains an internal model of the current situation and updates that model the moment prediction fails substantially.[10] When an action goal changes, space shifts, a new person appears, or the causal structure turns over, an event boundary forms, and these boundaries in turn shape the later units of memory.
Event boundaries give an important clue for thinking about the discreteness of consciousness. If major changes in conscious content occur not at fixed intervals but at the moments when the information structure changes, then the time of consciousness is organized by events rather than by clocks. These boundaries are not fixed. Watching the same video, a person learning a recipe and a person tracking a character's emotions treat different moments as important. Events are nested across multiple scales. Reaching out a hand is a short event, preparing a meal is a longer one, and the social context of serving that meal is longer still. This hierarchy and task dependence do not fit a fixed-frame hypothesis well. They support instead a picture in which quasi-stable states of several scales form on top of continuous neural processes, and states transition when prediction error and goal relevance grow large.
Computational models of the temporal structure of consciousness can be divided broadly into three. In a continuous model, content changes slightly at every moment along with input. In a fixed-frame model, the entire content is updated at regular intervals. In a hybrid model, underlying processing proceeds continuously, but stable content transitions into a new state when an event occurs. Bound together with the fewest assumptions, the current evidence favors the third model. Sensory input and neural state keep changing, attentional rhythm modulates the gain of particular information, recent information is integrated within a short temporal window, and when prediction error, salience, and goal relevance grow large enough, current content is reorganized into a new quasi-stable state. Computationally, this can be summarized as follows.
dz/dt = f(z(t), x(t))
c(k + 1) = B(z[t_k - tau : t_k])
when g(z(t), x(t), m(t)) > thetaHere z(t) is the continuously changing neural state, x(t) is sensory input, and m(t) is a modulatory variable such as attention, goal, or prediction error. When the function g crosses the threshold theta, information from the recent temporal window is organized into new content c(k + 1). What matters is that the update time t_k is not regular. The fixed-frame model is the special case where t_k = kT. In an event-based model, by contrast, a single state can hold for a long time in a calm, predictable scene, while a scene full of change can trigger several updates at short intervals. This model explains at once why consciousness looks frame-like and why it actually feels continuous. Content is stable for a certain period, so it is experienced as one scene, but underneath it sensory evidence and internal state keep changing. What is discrete is not the existence of the neural process but the reorganization of content. We cannot settle that the hybrid model is correct. But in that it can place known rhythmic selection, nonlinear access, temporal integration, and event segmentation within one structure, it is the most economical working hypothesis.
The first thing to do when bringing artificial intelligence into consciousness research is to draw a boundary. That a model compresses information, remembers, and distinguishes events does not mean it has subjective experience. There is no agreed experiment for determining phenomenal consciousness in current AI systems, and similarity of computational function does not guarantee identity of experience.[11] Even so, AI is useful for isolating and testing the computational components of consciousness theories. In the brains of humans and animals, selection, integration, memory, and action happen all at once, but in artificial systems each function can be implemented and removed separately. We can directly compare which bottlenecks are necessary, whether the update cycle must be fixed, and whether it is advantageous to remember only at event boundaries.
DeepMind's Perceiver iteratively compresses vast sensory input into a small latent array.[12] Rather than all input taking part in deep computation with equal weight, a limited latent space selectively takes information in. It is an architecture showing that having a bottleneck does not require input processing to be frame-based. The Allen Institute's MERLOT Reserve learns video, language, and sound together to infer temporally connected event structure.[13] Instead of storing the continuity of pixels and audio as-is, it forms representations useful for predicting what happened first and what will happen next. It is a computational case of a continuous stream being compressed into event-level internal representations. DeepMind's Differentiable Neural Computer and Neural Episodic Control show another separation of time.[14,15] One reads from and writes to external memory selectively; the other separates slowly changing network weights from rapidly updated episodic memory. It means that sensory processing, current state, event memory, and long-term learning can move at different speeds. In embodied-agent benchmarks such as ALFRED, past actions change the environment, and those changes become conditions for the next action.[16] The agent cannot finish the task by classifying the current screen alone; it must track, at the event level, what it has already picked up, which door is open, and which goal has been completed.
These models do not explain consciousness. But they turn the computations that consciousness theories implicitly demand into explicit parts. The most direct experiment would compare three agents under the same data and compute budget. One updates its state at every moment, one updates only at a fixed interval, and one updates when prediction error or goal relevance crosses a threshold. Measuring long-term prediction, event memory, action success rate, and computational cost together would show which temporal structure is adaptively advantageous. Such experiments do not answer whether AI is conscious. They answer the more restricted and testable question of what temporal structure is needed to build functions that resemble consciousness.
The continuous model, the fixed-frame model, and the event-based hybrid model can all explain many current results after the fact. To tell them apart we need experiments in which each model can fail. We would have to measure, within the same trials, whether early sensory representations change smoothly with stimulus strength while working memory and reportability switch abruptly at a particular moment; and to see whether updating aligns with the clock or with prediction error, we would check whether, over the same elapsed time, predictable scenes produce few state transitions while scenes with changing causal structure produce many. We would compare conditions that require report with conditions that do not, to see whether signals related to conscious content are really appearing because of report preparation and motor planning; and instead of only hunting for periodicity in behavioral performance, we would align the phase of a particular rhythm to stimulus timing, or perturb the rhythm with noninvasive stimulation, and see whether the temporal structure of perception changes as predicted. Above all, after fitting the three models to the same data, we would compare how accurately each predicts state-transition timing and reports in unseen trials. Consciousness research has no shortage of plausible narratives. What separates the models is prediction on new data.
The current evidence does not support the idea that consciousness is a film replayed at a fixed speed. Nor has any universal clock been found that binds every sensation and thought to a single frame rate. Conversely, it is also hard to say that every aspect of consciousness changes smoothly. Perceptual sensitivity rides a rhythm, access has a threshold, content is constructed over time, and it is rapidly reorganized at event boundaries. So the question of whether consciousness is discrete must be answered level by level. Neural dynamics are largely continuous. Information selection can be periodically modulated. Access to report and working memory can be nonlinear. Conscious content can hold quasi-stably and then update according to events. But whether experience itself vanishes completely in between is not yet known.
The most conservative conclusion is this. Consciousness is less a sequence of fixed frames than a system in which flowing neural processes are organized into one content state for a short time and cross into a new state before an important event. From this perspective, a moment is not the smallest particle of consciousness. It is the way a continuous process briefly takes on a stable form.
References
- Block, N. (1995). On a Confusion about a Function of Consciousness. Behavioral and Brain Sciences, 18(2), 227-247. doi:10.1017/S0140525X00038188
- VanRullen, R. (2016). Perceptual Cycles. Trends in Cognitive Sciences, 20(10), 723-735. doi:10.1016/j.tics.2016.07.006
- Landau, A. N., & Fries, P. (2012). Attention Samples Stimuli Rhythmically. Current Biology, 22(11), 1000-1004. doi:10.1016/j.cub.2012.03.054
- Busch, N. A., Dubois, J., & VanRullen, R. (2009). The Phase of Ongoing EEG Oscillations Predicts Visual Perception. Journal of Neuroscience, 29(24), 7869-7876. doi:10.1523/JNEUROSCI.0113-09.2009
- Fiebelkorn, I. C. et al. (2018). A Dynamic Interplay within the Frontoparietal Network Underlies Rhythmic Spatial Attention. Neuron, 99(4), 842-853.e8. doi:10.1016/j.neuron.2018.07.038
- Del Cul, A., Baillet, S., & Dehaene, S. (2007). Brain Dynamics Underlying the Nonlinear Threshold for Access to Consciousness. PLOS Biology, 5(10), e260. doi:10.1371/journal.pbio.0050260
- Cogitate Consortium et al. (2025). Adversarial Testing of Global Neuronal Workspace and Integrated Information Theories of Consciousness. Nature, 642, 133-142. doi:10.1038/s41586-025-08888-1
- Herzog, M. H., Kammer, T., & Scharnowski, F. (2016). Time Slices: What Is the Duration of a Percept? PLOS Biology, 14(4), e1002433. doi:10.1371/journal.pbio.1002433
- Fekete, T., Van de Cruys, S., Ekroll, V., & van Leeuwen, C. (2018). In the Interest of Saving Time: A Critique of Discrete Perception. Neuroscience of Consciousness, 2018(1), niy003. doi:10.1093/nc/niy003
- Zacks, J. M. et al. (2007). Event Perception: A Mind-Brain Perspective. Psychological Bulletin, 133(2), 273-293. doi:10.1037/0033-2909.133.2.273
- Butlin, P. et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.
- Jaegle, A. et al. (2021). Perceiver: General Perception with Iterative Attention. Proceedings of ICML 2021, PMLR 139, 4651-4664.
- Zellers, R. et al. (2022). MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound. Proceedings of CVPR 2022.
- Graves, A. et al. (2016). Hybrid Computing Using a Neural Network with Dynamic External Memory. Nature, 538, 471-476. doi:10.1038/nature20101
- Pritzel, A. et al. (2017). Neural Episodic Control. Proceedings of ICML 2017, PMLR 70, 2827-2836.
- Shridhar, M. et al. (2020). ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. Proceedings of CVPR 2020, 10740-10749.