When we started working on moment detection, we needed to understand what the system should be optimizing for. The obvious answer -- "find the good moments" -- is not an answer a signal detection model can use. We needed to identify what measurable properties of a clip correlate with sharing behavior. What we found from analyzing our early-access pilot exports is a set of recurring signal patterns that appear across content types and platform destinations.
Audio energy and the peak signature
The most consistent signal in shareable clips is an audio energy pattern that we call the peak signature: a transition from near-baseline audio to an elevated peak, followed by a sustained elevated period. The transition itself -- the moment where the energy jumps -- is almost always present in the first 15 seconds of a clip that gets shared. Clips that open on an already-elevated audio level without that transition tend to share less often, even when the underlying content quality is similar.
This makes intuitive sense once you think about how clips are consumed in a feed. The initial moments of a clip are competing for attention against all the other content a viewer is scrolling through. An energy transition in the opening window creates the audio equivalent of a visual hook -- something changes, and the viewer's attention follows the change. Clips that open on a static audio level require the viewer to have already committed their attention before the interesting part begins.
Face visibility and expression
In content types where the creator's face is visible, face-visibility windows correlated with elevated audio energy are a strong predictor of sharing. Clips where the face is obscured, off-camera, or low-visibility during the peak audio window perform worse on share metrics than clips where the face is clearly present. This holds even when the audio content is interesting -- the combination of an audible reaction and a visible face performing that reaction produces a more shareable unit than either element alone.
The expression specificity matters. Visible genuine surprise, laughter, or strong opinion delivery all correlate with sharing. Neutral face with elevated audio -- for example, a creator saying something important in a calm, level tone -- shows weaker sharing patterns than the same statement delivered with visible emotion. Shareable clips often feel like they could be used as a reaction meme even without audio context. If the visual tells a story independently of the words, the clip travels well.
Duration and the completion signal
Clips that achieve high completion rates (viewers watch to the end) correlate strongly with sharing behavior across all three platforms in our pilot data. The completion signal tells the algorithm that the clip held attention, which in turn drives wider distribution. The duration at which completion rates drop is content-dependent, but our pilot exports showed a consistent pattern: clips in the 25 to 45 second range maintained higher completion rates than clips in the 50 to 75 second range, even when the longer clips were rated as "better content" by the creator.
This is not an argument for always making shorter clips. It is an observation that length needs to be calibrated against the pace and density of the content. A 55-second clip where every second carries content that earns the viewer's attention is different from a 55-second clip that has 15 seconds of strong content and 40 seconds of setup and denouement. The detectable version of this pattern is speech pace density: how much information per second is in the audio in the first and last quarters of the clip relative to the middle.
What the signals do not capture
Signal patterns are predictors, not guarantors. The strongest audio energy peak in a stream may be the creator coughing. A face-visible window with high expression visibility may be a moment of genuine frustration that the creator would not want posted. The clips that go most widely viral are often the ones that contain a genuinely novel element -- a specific game event, a one-in-a-million coincidence, a uniquely phrased opinion -- that no signal pattern could have predicted. Detection finds the candidates; the creator still makes the call.