.NET / SQL / Enterprise Engineering

Psychoacoustics of Deep Auditory Absorption: A Research Review for Generative Web Environments

Report summary

The development of browser-based generative sound environments, such as the proposed standalone Sound Studio and the Advanced Sound Studio panel within the Visualization Studio for mesmerization.com, necessitates a rigorous, evidence-based understanding of human auditory perception, spatial immersio

Status
Research archive item
Category
.NET / SQL / Enterprise Engineering
Length
4,794 words
Reading time
22 minutes
Report type
evaluation

Key topics

  • .NET / SQL / Enterprise Engineering
  • .NET
  • SQL
  • Enterprise Engineering
  • AI
  • Rust
  • Physics
  • Research Archive
  • Strategy

Research provenance

Archive status
Research archive item
Content identity
sha256:7b3acdfeed1e13c422131b81ab46c37519984c1b56847fab0d2e8f0e6138e607

For citation, use the report title and canonical URL. Archival presence does not establish authorship or promote report statements into portfolio evidence.

This page renders the archived Markdown as safe, formatted HTML. It is background research and does not become a portfolio claim without evidence review.

Full report

On this page

Introduction

The development of browser-based generative sound environments, such as the proposed standalone Sound Studio and the Advanced Sound Studio panel within the Visualization Studio for mesmerization.com, necessitates a rigorous, evidence-based understanding of human auditory perception, spatial immersion, and sustained attention. The primary objective of this acoustic architecture is to foster a state of deep auditory absorption—often described subjectively as perceptual immersion, fascination, or mesmerization—without relying on visual stimuli. To achieve this effectively and safely, the underlying audio engine must be grounded in established psychoacoustic principles and neurophysiological realities rather than speculative, pseudo-scientific claims.

It is a foundational premise of this analysis that browser-based audio cannot force a listener into a state of clinical hypnosis, exert mind control, or medically treat specific neurological, psychiatric, or psychological conditions. A prominent example of such contested claims is the theory of direct brainwave entrainment via binaural beats. Proponents argue that presenting two slightly different frequencies to each ear forces macroscopic electroencephalographic (EEG) oscillations to synchronize with the difference frequency (e.g., forcing the brain into a theta or delta state to guarantee relaxation)1. However, comprehensive meta-analyses and rigorous physiological studies consistently demonstrate that while binaural beats do create a perceptual illusion of modulation at the superior olivary complex, they fail to induce the widespread cortical EEG entrainment, cognitive enhancement, or emotional arousal often claimed by commercial entities3. Although monaural beats exhibit stronger physical amplitude modulation and elicit more salient auditory steady-state responses than binaural beats, neither acts as a guaranteed neurological switch6. Therefore, a scientifically valid approach to auditory fascination must discard these speculative shortcuts.

Instead, continuous sound becomes compelling through the nuanced manipulation of how the human brain parses, groups, and attends to complex acoustic information over time. This report systematically examines the psychoacoustics of auditory scene analysis, expectation formation, and altered subjective time perception. By distinguishing useful artistic phenomena from risky shortcuts—such as extreme volume, excessive bass, or rapid high-depth amplitude modulation—this review establishes a deterministic digital signal processing (DSP) framework for a Web Audio API-driven generative environment that feels continuous, infinite, alive, spacious, and deeply absorbing while remaining entirely intelligible and comfortable.

Evidence-Ranked Synthesis of Auditory Phenomena

To design a deterministic generative sound engine, auditory phenomena must be strictly categorized by their empirical reliability, physiological validity, and utility in sustaining non-fatiguing attention. The following synthesis ranks relevant research concepts from the most universally established psychophysical laws to the highly contested claims that must be excluded from the design architecture.

The first tier consists of established psychophysical and psychoacoustic principles, which represent fundamental mechanical and neurological functions of the human auditory system. These phenomena are highly predictable, mathematically modelable, and universal across normal-hearing populations. A critical component of this tier is the understanding of equal-loudness contours, formalized in ISO 226:2023. The human auditory system does not perceive all frequencies at an equal amplitude; sensitivity peaks significantly in the 2–5 kHz range and drops precipitously at lower (sub-bass) and higher frequencies8. Designing for low-volume immersion requires compensating for these contours without introducing excessive low-frequency energy that could cause upward spread of masking, which would obscure mid-range intelligibility8. Furthermore, fundamental phenomena such as spectral and temporal masking dictate how simultaneous and sequential sounds occlude one another, while concepts like roughness and fluctuation strength define the strict boundaries between biological comfort and acoustic harshness10.

The second tier encompasses cognitive, attentional, and contextual phenomena. These rely on higher-order cognitive processing, working memory, and the top-down modulation of auditory scene analysis. While these mechanisms are robustly supported by literature, they exhibit slightly more inter-individual variability than basic mechanical acoustics. This includes the auditory continuity illusion (temporal induction), wherein the brain interpolates a missing sound if an interruption is filled with a sufficiently loud and spectrally broad noise, maintaining the perception of an unbroken soundscape12. It also includes the Attentional Gate Model of time perception, which explains how deep sensory absorption diverts cognitive resources away from internal timekeeping, leading to subjective time compression14. Auditory salience, driven by the brain's continuous statistical modeling of the environment, dictates how deviations in pitch, timbre, or rhythm trigger attentional capture16.

The third tier consists of contested, speculative, or pseudoscientific claims, which lack robust peer-reviewed validation and pose a risk to the credibility of a generative audio environment. As previously noted, the assertion that binaural beats can reliably force specific EEG states for therapeutic outcomes falls into this category3. Similarly, the concept of subliminal messaging for neurological change—the idea that inaudible affirmations embedded beneath noise can bypass the conscious mind to program neuroplasticity—lacks empirical support for long-term cognitive alteration18. The generative engine must strictly avoid attempting to leverage these disproven concepts, focusing entirely on the demonstrable mechanics of acoustics and perception.

Taxonomy of Auditory Absorption Mechanisms

Deep auditory absorption occurs when the auditory system is sufficiently engaged to prevent habituation, yet not so aggressively stimulated that it triggers fatigue, tension, or a stress response. The environment must feel alive and mysterious. This delicate balance is achieved through the orchestration of specific perceptual mechanisms, detailed below.

Auditory Scene Analysis, Stream Segregation, and Object Formation

The auditory system does not process the world as a raw, monolithic waveform; rather, it organizes acoustic energy into discrete "auditory objects" defined primarily by pitch, location, spectral envelope, and time19. According to Kubovy and Van Valkenburg, auditory object formation operates through parallel "what" and "where" processing streams, analogous to visual perception, allowing listeners to track the identity and location of a sound source simultaneously21. In a generative environment, maintaining auditory fascination requires playing continuously with the boundary between stream fusion (perceiving sounds as a single cohesive object) and stream segregation (perceiving sounds as multiple distinct objects).

Using the classic ABA tone paradigm (where A and B are tones of differing frequencies), van Noorden demonstrated that when the frequency separation between tones is small and the repetition rate is slow, the tones fuse into a single rhythm, an effect known as temporal coherence22. As the frequency separation increases or the tempo accelerates, the percept splits, or fissions, into two distinct, parallel streams. By subtly morphing the spectral centroid or the temporal spacing of generative voices, a sound engine can cause a single sustained texture to organically pull apart into multiple distinct melodic voices, and then gradually coalesce back into one. This constant, threshold-level shifting forces the auditory system to continuously update its scene analysis, sustaining endogenous (top-down) attention without listener fatigue25.

Sustained tones versus continuously transforming timbres play a vital role here. A perfectly static sustained tone leads to rapid sensory adaptation; the neurons in the auditory cortex reduce their firing rate, and the sound effectively disappears from conscious awareness. Conversely, continuously transforming timbres—achieved through amplitude-envelope complexity, phase shifts, and slow filter modulations—keep the auditory system actively tracking microvariations, preventing full habituation while avoiding the startle responses associated with abrupt transients.

Perceptual Filling-In and Auditory Continuity

The auditory continuity illusion, also known as temporal induction, is a powerful hallmark of top-down perceptual organization. When a continuous sound, such as a drone or a slow vocal glide, is momentarily interrupted by a brief, spectrally overlapping noise (like a simulated wind gust or a wave crash), the listener's brain actively hallucinates the continuation of the tone through the noise12. The brain applies a "No Discontinuity Rule," assuming that the softer tone was merely masked by the louder noise rather than terminated12.

Neurologically, this phenomenon is evidenced by a lack of mismatch negativity (MMN) and sustained theta-band phase-locking in the auditory cortex, indicating that the brain treats the sound as physically uninterrupted27. For an application focused on mesmerization, utilizing the continuity illusion is an invaluable artistic tool. It allows the generative engine to introduce silence, micro-pauses, or masking textures that give the ear and the transducer a brief physical rest, while the listener's brain perceives an infinite, unbroken, and majestic soundscape12. This prevents repetitive annoyance and auditory fatigue while enhancing the mysterious, expansive quality of the audio.

Spectral Masking, Temporal Masking, and Comodulation

Masking occurs when the threshold of audibility for one sound is raised by the presence of another. Spectral masking dictates that a loud noise will obscure a quieter tone within the same critical band. Temporal masking occurs in the time domain: a loud transient can mask softer sounds that immediately follow it (forward masking) and, remarkably, softer sounds that immediately precede it (backward masking). Relying on heavy temporal masking via startling transients is an anti-pattern for a relaxing environment, as it obliterates the subtle microvariations of the soundscape.

However, masking can be manipulated to create profound perceptual depth through Comodulation Masking Release (CMR). Hall discovered that if a masking noise is amplitude-modulated coherently across multiple, widely separated frequency bands (comodulation), the auditory system uses the synchronized dips in the noise envelope to "listen in" and unmask a target sound30. CMR demonstrates that the brain actively groups sounds that share a common amplitude envelope, a principle known as common fate32. In a generative environment, applying a highly correlated, slow low-frequency oscillator (LFO) to the amplitude of both a broad background texture and a faint foreground melody will suddenly cause the melody to emerge perceptually from the noise floor. This creates a profound sense of acoustic mystery; the sound field feels "alive" because the environment is breathing as a single cohesive, organic entity.

Expectation, Salience, and Predictability Gradients

Sustaining auditory fascination requires navigating the precise boundary between repetition and novelty. Auditory salience is driven by Spectro-Temporal Receptive Fields (STRFs) in the primary auditory cortex, which dynamically adapt to the statistical regularities of the acoustic environment17. If a generative loop relies on exact repetition, the STRFs quickly adapt, neural responses attenuate, and the listener experiences habituation. If the environment is entirely stochastic and unpredictable, it creates high cognitive load, violating expectation formation and causing irritation.

Absorption thrives on a predictability gradient—a delicate balance where a schema is established and then gently violated. This involves managing endogenous versus exogenous auditory attention. Exogenous attention is stimulus-driven and bottom-up; a sudden, loud transient will forcibly capture exogenous attention, triggering an orienting reflex or a startle response. Endogenous attention is goal-directed and top-down; it is the state of a listener willingly choosing to explore the depths of a soundscape. The generative engine must use continuous microvariation, subtle rhythmic grouping, and shifting modulation spectra to gently capture exogenous attention and smoothly transition the listener into a state of sustained endogenous attention. By slowly transforming timbres rather than abruptly introducing new elements, the system respects the listener's expectation formation while providing enough novelty to prevent sensory adaptation.

Altered Subjective Time Perception and Auditory Looming

A successful generative environment alters the listener's subjective experience of time. The Attentional Gate Model (AGM) of time perception, proposed by Zakay and Block, posits that an internal pacemaker generates continuous temporal pulses that are collected by an accumulator14. An "attentional gate" controls the flow of these pulses. When a listener is bored, stressed, or exposed to highly threatening sounds, attention is directed toward time monitoring; the gate opens fully, pulses accumulate rapidly in working memory, and time is perceived to drag or dilate35.

Conversely, deep auditory fascination redirects cognitive resources away from the internal temporal pacemaker and entirely onto the sensory experience. The attentional gate partially closes, fewer temporal pulses accumulate, and the listener experiences temporal compression—the profound feeling that hours have passed in mere minutes14.

To maintain a closed attentional gate, the system must carefully manage the auditory looming bias. Evolutionary biology has created a perceptual asymmetry wherein approaching sounds (increasing in intensity) are perceived as more salient, significantly closer, and faster-moving than receding sounds (decreasing in intensity)35. While a very slow looming effect can gently guide attention and build subtle anticipation, aggressive looming triggers the brainstem's threat evaluation response, flooding the system with cortisol, violently opening the attentional gate, and destroying the state of deep absorption.

Properties of the Immersive Sound Field

To ensure the standalone Sound Studio feels spacious, alive, and comfortable without relying on visualizations, the DSP architecture must prioritize spatial depth, organic timbral modulation, and a strict adherence to spectral balance.

Spatial Immersion, Perceived Distance, and Reverberant Depth

Perceived distance, spatial immersion, and perceived motion are governed largely by the Direct-to-Reverberant Ratio (DRR) and the Interaural Cross-Correlation Coefficient (IACC)38. A high DRR, consisting mostly of direct sound, places an auditory object intimately close to the listener's head. While initially engaging, high-DRR sounds can become claustrophobic and fatiguing over long periods. Conversely, a low DRR pushes the sound source back into a reverberant field, creating vastness, environmental context, and reverberant depth40.

Furthermore, IACC measures the similarity of the acoustic signals arriving at the left and right ears. A high IACC indicates a narrow, centered sound, while a low IACC, utilizing highly decorrelated signals, creates a powerful sense of spatial envelopment and wideness. Dynamically modulating the IACC and DRR of different generative streams can simulate perceived motion, making the soundscape feel as though it is organically expanding and contracting.

Timbral Dynamics: Consonance, Dissonance, Roughness, and Fluctuation Strength

The emotional and physical response to continuous sound is heavily dictated by amplitude-envelope complexity and modulation spectra. According to the foundational psychoacoustic models of Zwicker and Fastl, the rate of amplitude modulation dictates whether a sound is perceived as relaxing or distressing10.

Modulation rates around 4 Hz produce a sensation known as fluctuation strength. This frequency mimics the cadence of human walking, a resting heartbeat, and ocean waves. Maximizing fluctuation strength induces a sense of rhythmic grouping, biological synchrony, and calm41. Conversely, modulation rates between 15 Hz and 150 Hz, peaking around 70 Hz, produce timbral roughness. Roughness is perceived as buzzing, harshness, and severe dissonance, structurally similar to an animal's growl or a revving engine10.

In the context of tension and release, the generative engine may occasionally allow two frequencies to drift into the same critical band to create brief dissonance and slight roughness, building aesthetic tension. However, this must be strictly temporary, smoothly resolving back into consonance and low-frequency fluctuation strength to provide psychological release. Sustained high-depth amplitude modulation in the roughness range is a risky shortcut that must be prohibited, as it physically fatigues the basilar membrane.

Spectral Balance and Auditory Fatigue

The spectral centroid is the mathematically defined "center of gravity" of a sound's frequency spectrum, heavily correlated with the perceptual attribute of "brightness"44. Sustained exposure to continuous sounds with a high spectral centroid (excessive high-frequency energy) leads to rapid auditory fatigue, potential exacerbation of tinnitus, and an involuntary withdrawal of endogenous attention47. To maintain comfort over long sessions, the system must utilize dynamic low-pass filters to keep the overall spectral centroid centered in the lower-mid frequencies, deploying high-frequency shimmer only as a transient, rare reward. Similarly, extreme low-frequency energy (excessive sub-bass) must be avoided, as it triggers the acoustic reflex in the middle ear and muddies the generative composition.

Practical Browser-Audio Design Principles and Parameter Bounding

Implementing these rigorous psychoacoustic principles within a browser environment requires navigating the specific constraints, capabilities, and performance bottlenecks of the Web Audio API.

Historically, complex browser audio relied on the ScriptProcessorNode, which executed on the main JavaScript thread. This legacy approach was highly susceptible to garbage collection pauses, dropped frames, and severe latency49. The modern standard for deterministic DSP is the AudioWorklet, which runs custom C++ or Rust code compiled to WebAssembly (WASM), or deterministic JavaScript, on a dedicated, high-priority audio rendering thread50.

The Web Audio API processes audio in deterministic render quanta of exactly 128 samples50. To maintain the delicate, sample-accurate phase relationships required for spatial immersion (IACC) and comodulation (CMR), the generative engine must be built entirely within AudioWorklets. Parameter automations should utilize methods such as AudioParam.linearRampToValueAtTime mapped against the precise BaseAudioContext.currentTime to guarantee flawless microvariations53.

Continuous generative audio must avoid the "loudness wars" and target a comfortable, standard broadcasting level. Utilizing the ITU-R BS.1770-4 standard and EBU R128 guidelines, the continuous integrated loudness of the environment should be bounded between \-24 LUFS and \-16 LUFS54. Because users frequently listen to generative environments at low volumes, the system should gently implement a dynamic inverse ISO 226:2023 equalization curve. At low sound pressure levels, the human ear is dramatically less sensitive to bass and treble; a subtle dynamic EQ that boosts these extremes when the overall LUFS is low ensures the environment retains its perceived depth without requiring the user to dangerously raise the volume of their headphones8.

The following anti-patterns and risky shortcuts must be strictly prohibited in the generative logic:

  • Extreme Volume and Piercing Resonances: High-Q resonant peaks in the 2–5 kHz range will cause immediate pain and permanent fatigue due to the ear's natural resonance.
  • Startling Transients: Sharp attack times trigger the brainstem's startle response, flooding the system with adrenaline and destroying the closed attentional gate35.
  • Excessive Bass: Sustained frequencies below 30 Hz mask intricate mid-range details and can distort consumer transducer equipment.
  • Monotonous Loops: Exact sample-for-sample repetition causes STRF habituation within minutes, rendering the sound functionally invisible to the auditory cortex16.
  • Rapid High-Depth Amplitude Modulation: Modulating volumes at 20–100 Hz introduces severe roughness, causing psychological distress and physical fatigue10.

Testable Hypotheses and Physical Listening Tests

To refine the Sound Studio empirically, the following hypotheses can be tested during generative deployment:

1. Environments utilizing Comodulation Masking Release—where background noise and foreground drones share an identical, synchronized 0.1 Hz amplitude envelope—will yield a significantly longer average session time than environments with uncorrelated envelopes, due to increased perceptual cohesion and common fate grouping.

2. Algorithmically lowering the global spectral centroid by 10% every 15 minutes will delay the onset of self-reported auditory fatigue and result in higher subjective relaxation scores compared to a static, brighter frequency spectrum.

3. Deploying the auditory continuity illusion by introducing 2-second gaps in a drone, filled with spectrally matched pink noise, will be perceived as a continuous, unbroken soundscape by users, reducing habituation without breaking the hypnotic state.

While DSP measurements provide mathematical boundaries, physical listening tests are irreplaceable. A mathematically perfect low-frequency wave generated in the Web Audio API can cause Intermodulation Distortion (IMD) on the physical diaphragm of a consumer laptop speaker, generating high-frequency harmonic artifacts that the DSP cannot measure. Furthermore, spatial audio processed through generic binaural HRTFs can sound vastly enveloping to one listener and severely colored to another depending on their unique ear geometry. Finally, long-term attentional fatigue cannot be calculated by an algorithm; only multi-hour physical listening sessions can determine if the harmonic structure of a patch induces the sensation of time-dragging.

Mapping Research Concepts to User-Facing Controls

To provide users with meaningful control over the generative environment without requiring expertise in acoustics, complex psychoacoustic parameters should be abstracted into intuitive, evocative sliders.

Psychoacoustic ConceptDeterministic DSP MechanismProposed User-Facing ControlEffect on Auditory Absorption
Interaural Cross-Correlation (IACC)Mid/Side processing; delaying the Side channel by 1–15 ms; decorrelating left/right phase."Vastness" or "Width"High width increases spatial envelopment, making the listener feel placed inside the environment rather than observing it from outside.
Direct-to-Reverberant Ratio (DRR)Adjusting the ratio of the dry signal to a highly diffuse algorithmic convolution reverb."Distance" or "Depth"Pushes the sound source away from the skull, reducing claustrophobia and increasing the sense of a mysterious, infinite reverberant space.
Spectral Centroid / ISO 226:2023Dynamic low-pass filter cutoff combined with a subtle high-shelf attenuation."Warmth" or "Softness"Lowers high-frequency fatigue. Increases listening comfort for multi-hour sessions while maintaining necessary mid-range intelligibility.
Fluctuation Strength (Zwicker)Global amplitude and filter cutoff modulation via a synchronized LFO set between 0.1 Hz and 4 Hz."Breath" or "Pulse"Mimics biological rhythms. Enhances the "aliveness" of the soundscape, anchoring endogenous attention in a hypnotic, predictable cadence.
Stream Segregation (ASA)Modulating the pitch separation and rhythmic spacing of generative arpeggios."Complexity" or "Fracture"Forces a single drone to perceptually split into multiple voices, preventing habituation by gently demanding continuous cortical re-analysis.

Deterministic DSP Measurements as Proxies

To maintain policy compliance and ensure the generative engine remains within safe, absorbing parameters, the following deterministic DSP measurements must be continuously calculated within the AudioWorklet.

 

Target Psychoacoustic StateDSP Measurement ProxySafe Bounded Range
Non-Fatiguing LoudnessIntegrated LUFS (ITU-R BS.1770-4) using a 400ms sliding window.\-24 LUFS to \-16 LUFS55.
Absence of Harsh RoughnessFast Fourier Transform (FFT) analysis of amplitude modulation sidebands.Modulation peaks must remain below 10 Hz. Strong programmatic suppression of 20–150 Hz AM depth10.
Comfortable BrightnessReal-time Spectral Centroid calculation.Centroid moving average bounded below 2.5 kHz for continuous drones; occasional transient peaks strictly limited47.
Predictability GradientSpectral Flux (Euclidean distance between normalized spectra of successive frames).Must maintain a non-zero, low-variance baseline to ensure continuous microvariation without startling spectral jumps60.
Spatial EnvelopmentCross-correlation function of Left and Right output buffers.IACC [Figure omitted from source export] for background textures to ensure maximum perceived wideness and immersion without phase cancellation.

Proposed Auditory Absorption Policy

To ensure that the mesmerization.com generative sound engine safely and effectively induces deep auditory absorption without crossing into pseudoscience, acoustic fatigue, or browser instability, the following policy is proposed for the deterministic Web Audio engine.

The platform shall maintain a strict demarcation of claims, explicitly avoiding all clinical, medical, or pseudo-neurological terminology. No claims shall be made regarding the enforcement of hypnosis, the treatment of psychological conditions, or the direct entrainment of specific EEG brainwave states via binaural or monaural beats. The system's efficacy is to be defined strictly through the lens of aesthetic fascination, perceptual immersion, and the facilitation of sustained endogenous attention through recognized psychoacoustic principles.

Architectural determinism is mandated. All core generative DSP, including spatialization, masking release, and equal-loudness contouring, must be executed within isolated AudioWorklet threads to guarantee 128-sample low-latency determinism, rendering the audio immune to main-thread UI garbage collection and ensuring phase accuracy.

Global integrated loudness must be strictly and algorithmically constrained between \-24 LUFS and \-16 LUFS. Peak limiting must be applied invisibly ahead of the output destination to prevent any transient spikes from triggering a startle response or opening the listener's attentional gate.

To prevent the induction of sensory roughness and physiological stress, amplitude and filter modulation rates must be strictly bounded. Modulations shall operate exclusively in the fluctuation strength domain (0.1 Hz to 8 Hz), with a hard programmatic limit prohibiting sustained modulation depths in the roughness domain (15 Hz to 150 Hz). Occasional crossings into this range are permitted only briefly for the purpose of artistic tension and release.

To eliminate high-frequency auditory fatigue during multi-hour sessions, generative voices must be actively monitored for spectral balance. If the global spectral centroid exceeds 3.0 kHz for a sustained period, the engine must autonomously lower the cutoff frequencies of global low-pass filters to restore timbral warmth.

Finally, the generative engine is prohibited from deploying exact, bit-for-bit looping buffers. To prevent STRF habituation and maintain a compelling predictability gradient, all textural layers must incorporate continuous, slow-moving microvariations in phase, pitch, and amplitude envelope complexity, ensuring the soundscape remains perceptually infinite, alive, and profoundly absorbing.

Works cited

1. (PDF) Binaural beats to entrain the brain? A systematic review of the, https://www.researchgate.net/publication/370900678\_Binaural\_beats\_to\_entrain\_the\_brain\_A\_systematic\_review\_of\_the\_effects\_of\_binaural\_beat\_stimulation\_on\_brain\_oscillatory\_activity\_and\_the\_implications\_for\_psychological\_research\_and\_intervention

2. Binaural Beat: A Failure to Enhance EEG Power and Emotional, https://www.researchgate.net/publication/321078055\_Binaural\_Beat\_A\_Failure\_to\_Enhance\_EEG\_Power\_and\_Emotional\_Arousal

3. Binaural Beat: A Failure to Enhance EEG Power and Emotional, https://www.frontiersin.org/journals/human-neuroscience/articles/10.3389/fnhum.2017.00557/full

4. Binaural Beat: A Failure to Enhance EEG Power and Emotional, https://pmc.ncbi.nlm.nih.gov/articles/PMC5694826/

5. Binaural Beat: A Failure to Enhance EEG Power and Emotional, https://pubmed.ncbi.nlm.nih.gov/29187819/

6. music with monaural beats reduces anxiety and improves mood in a, https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2025.1539823/pdf

7. Binaural Beats through the Auditory Pathway: From Brainstem to, https://pmc.ncbi.nlm.nih.gov/articles/PMC7082494/

8. Engineering Noise Control \[6 ed.\] 2023001141, 9780367414788, https://dokumen.pub/engineering-noise-control-6nbsped-2023001141-9780367414788-9780367414795-9780367814908.html

9. Smiley face curve \- Grokipedia, https://grokipedia.com/page/Smiley\_face\_curve

10. Beating and Roughness \- Wolfram Demonstrations Project, https://demonstrations.wolfram.com/BeatingAndRoughness/

11. ECMA-418-2, 2nd edition, December 2022, https://ecma-international.org/wp-content/uploads/ECMA-418-2\_2nd\_edition\_december\_2022.pdf

12. Dynamics of the Auditory Continuity Illusion \- PMC \- NIH, https://pmc.ncbi.nlm.nih.gov/articles/PMC8217826/

13. Confidence in auditory perceptual completion \- Oxford Academic, https://academic.oup.com/nc/article/2025/1/niaf018/8236441

14. Time perception, attention, and memory: A selective review, https://www.researchgate.net/publication/259454874\_Time\_perception\_attention\_and\_memory\_A\_selective\_review

15. How Self-Paced Visual Exploration Compresses Subjective Time, https://www.biorxiv.org/content/10.64898/2026.07.02.734699v1.full.pdf

16. Pitch, timbre and intensity interdependently modulate neural ... \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC7363547/

17. Hearing Research, https://hearingbrain.org/pdf/fritz\_et\_al\_hearingres2007.pdf

18. How To Speed Up Subliminal Results \- Make Subliminals Work Faster, https://alteredmindwaves.com/how-to-speed-up-subliminal-results-make-subliminals-work-faster/

19. Binding of verbal and spatial features in auditory working memory, https://doi.org/10.1016%2Fj.jml.2009.03.001

20. Auditory Perception \- Stanford Encyclopedia of Philosophy, https://plato.stanford.edu/archives/fall2009/entries/perception-auditory/

21. Auditory and visual objects \- TIMARA, https://timara.net/s/KubovyVanValkenburg2001.pdf

22. (PDF) Temporal predictability as a grouping cue in the perception of, https://www.researchgate.net/publication/249996086\_Temporal\_predictability\_as\_a\_grouping\_cue\_in\_the\_perception\_of\_auditory\_streams

23. The role of spectral and periodicity cues in auditory stream, https://pubs.aip.org/asa/jasa/article-pdf/106/2/938/9366485/938\_1\_online.pdf

24. Auditory stream segregation in cochlear implant listeners, https://pubs.aip.org/asa/jasa/article-pdf/126/4/1975/15288299/1975\_1\_online.pdf

25. The Build-up of Auditory Stream Segregation: A Different Perspective, https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2012.00461/full

26. Auditory stream segregation of iterated rippled noises by normal, https://pmc.ncbi.nlm.nih.gov/articles/PMC5785299/

27. Neural Underpinnings of the Continuity Illusion in Musicians \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC12831042/

28. (PDF) The Neurophysiological Basis of the Auditory Continuity Illusion, https://www.researchgate.net/publication/5269910\_The\_Neurophysiological\_Basis\_of\_the\_Auditory\_Continuity\_Illusion\_A\_Mismatch\_Negativity\_Study

29. Neural mechanisms for illusory filling-in of degraded speech \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC2653101/

30. The effect of the preceding masking noise on monaural and binaural, https://www.biorxiv.org/content/10.1101/2021.11.05.467091v3.full

31. Comodulation masking release in the inferior colliculus by combined, https://pmc.ncbi.nlm.nih.gov/articles/PMC5336552/

32. Factors contributing to comodulation masking release with dichotic, https://pmc.ncbi.nlm.nih.gov/articles/PMC2600623/

33. Does Attention Play a Role in Dynamic Receptive Field Adaptation, https://pmc.ncbi.nlm.nih.gov/articles/PMC2077083/

34. Psychological time as information: the case of boredom \- Frontiers, https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2014.00917/full

35. Psychological and Neural Mechanisms of Subjective Time Dilation, https://pmc.ncbi.nlm.nih.gov/articles/PMC3085178/

36. The influence of threat on time perception in anxiety \- ResearchGate, https://www.researchgate.net/publication/232702505\_When\_time\_slows\_down\_The\_influence\_of\_threat\_on\_time\_perception\_in\_anxiety

37. No representational momentum for auditory looming stimuli \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC12913253/

38. Brain activity discriminates acoustic simulations of the same, https://www.biorxiv.org/content/10.1101/2024.08.29.610373v3.full-text

39. Google Sports Data, https://support.google.com/knowledgepanel/answer/9787176

40. Auditory Distance Perception in Wave Field Synthesis, https://projekter.aau.dk/projekter/files/852115960/MSc\_Thesis\_Emil\_Zawistowski.pdf

41. Roughness and Fluctuation Strength \- Ansys Help, https://ansyshelp.ansys.com/public/Views/Secured/corp/v261/en/Sound\_SAS\_UG/Sound/UG\_SAS/roughness\_and\_fluctuation\_strength\_132348.html

42. Modeling the fluctuation strength of technical sounds, https://pub.dega-akustik.de/DAGA\_2023/data/articles/000181.pdf

43. Fluctuation Strength (Vacil) Calculator \- MetricGate, https://metricgate.com/docs/fluctuation-strength-vacil/

44. A New Strategy for Control of Adaptive Digital Audio Effects, https://www.bosleymusic.com/Papers/DonBosleyThesis\_06172013\_Edit.pdf

45. Audio Bandwidth Extension \- Signal Processing Systems, https://www.sps.tue.nl/rmaarts/RMA\_papers/aar04pu6F.pdf

46. Assessment of Laying Hens' Thermal Comfort Using Sound ... \- PMC, https://pmc.ncbi.nlm.nih.gov/articles/PMC7013866/

47. Designing Sound, https://wp.ufpel.edu.br/labcomp/files/2022/09/PureData-Designing\_Sound-Andy\_Farnell.pdf

48. New Advances in Audio Signal Processing \- MDPI, https://mdpi-res.com/bookfiles/book/9283/New\_Advances\_in\_Audio\_Signal\_Processing.pdf?v=1782868131

49. Riding the latest waves on the web \- Superpowered, https://superpowered.com/riding-the-latest-waves-on-the-web

50. Web Audio API \+ AI: Why AudioWorklet \+ WASM Is the 2026, https://callsphere.ai/blog/vw9e-web-audio-api-ai-processing-audioworklet-2026

51. Frontend Architecture of a Voice Agent Interface \- Medium, https://medium.com/@ujjwaltiwari2/frontend-architecture-of-a-voice-agent-interface-6236bfc393ba

52. Web Audio API \- MDN Web Docs \- Mozilla, https://developer.mozilla.org/en-US/docs/Web/API/Web\_Audio\_API

53. Web Audio API 1.1 \- W3C, https://www.w3.org/TR/webaudio-1.1/

54. Beginner Music Production: Your Step-by-Step Guide \- SoundBridge, https://www.soundbridge.io/beginner-music-production-your-step-by-step-guide

55. A NEW APPROACH AND GUIDLINE FOR LOUDNESS IN GAME, https://his.diva-portal.org/smash/get/diva2:1790653/FULLTEXT01.pdf

56. ITU-R BS.1770-4 Loudness Measurement | PDF \- Scribd, https://www.scribd.com/document/612843562/R-REC-BS-1770-4-201510-I-MSW-E-1

57. EBU R 128 \- Wikipedia, https://en.wikipedia.org/wiki/EBU\_R\_128

58. Naturalistic Stimulus Structure Determines the Integration of, https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0070710

59. Loudness Control according to EBU R128 | En | Wiki.Audio, https://wiki.audio/En/0129

60. Sweet \[re\]production: Developing sound spatialization tools for, https://www.mcgill.ca/mpcl/files/mpcl/peters\_2010\_phdthesis.pdf