<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator><link href="https://nikhilr.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://nikhilr.io/" rel="alternate" type="text/html" /><updated>2026-07-27T22:13:02+08:00</updated><id>https://nikhilr.io/feed.xml</id><title type="html">Nikhil Raghavendra</title><subtitle>Studying intelligence, artificial and otherwise</subtitle><author><name>Nikhil Raghavendra</name></author><entry><title type="html">The Evolutionary Foundations of Mind Perception</title><link href="https://nikhilr.io/posts/Mind-Perception/" rel="alternate" type="text/html" title="The Evolutionary Foundations of Mind Perception" /><published>2025-12-25T00:00:00+08:00</published><updated>2025-12-25T00:00:00+08:00</updated><id>https://nikhilr.io/posts/Mind-Perception</id><content type="html" xml:base="https://nikhilr.io/posts/Mind-Perception/"><![CDATA[<p>As I scroll through social media, read the news, or watch YouTube, I constantly come across posts and articles about Artificial Intelligence (AI), AI agents, and even Artificial General Intelligence (AGI). These technologies are often framed as forces that could reshape society or threaten it if they are misaligned or wielded by bad actors. This flood of anxious commentary made me pause and ask a simpler question: why are we so afraid of a computer program we built, especially when it still cannot replicate the full spectrum of human intelligence and intellectual agency?</p>

<p>A realization followed. We are no longer reacting to ordinary software or to industrial robots that have worked in our factories for decades. Instead, we are confronting something deeper known as mind perception. This is the human capacity to perceive consciousness and to attribute internal mental life, subjective experience, and intentional agency to entities other than ourselves. Far from being a passive mirror of reality, this capacity is an active and evolutionarily constructed interface shaped by millions of years of selective pressure.</p>

<p>The core thesis in the literature is that mind perception is a survival adaptation rather than a tool designed to deliver philosophical truth. It is governed by Error Management Theory, which holds that the cost of failing to detect an agent in uncertain environments is far greater than the cost of mistakenly detecting one. This asymmetry fostered a Hyperactive Agency Detection Device and a teleological stance as we evolved, biasing human cognition toward false positives. We readily attribute a mind to wind-blown bushes, moving geometric shapes, and now, increasingly, to algorithms.</p>

<p>However, this detection system is not uniform. Once an entity is flagged as a potential agent, it is evaluated through secondary filters based on phylogenetic proximity and the moral dyad of agency and experience. We feel deep empathy for mammals that resemble us and share familiar social and affective cues, yet we struggle to recognize suffering in invertebrates or in entities that fall into the Uncanny Valley, where ambiguity can trigger ancient avoidance responses.</p>

<p>The rapid progression of AI and robotics marks a pivotal moment where these ancient Pleistocene adaptations for social cognition are being actively engaged and sometimes manipulated. This era witnesses the profound exploitation of deeply ingrained psychological mechanisms that were originally honed for navigating complex human social structures but are now being triggered by artificial entities.</p>

<p>One striking example is the amplification of the ELIZA effect, the unconscious tendency to attribute human-like characteristics or intelligence to a computer system. Modern Large Language Models (LLMs) use fluid and human-mimicking linguistic output to systematically activate the very cues that our brains associate with genuine agency, consciousness, and intent. This sophisticated natural language capability effectively capitalizes on our evolutionary predisposition to find a mind behind a voice.</p>

<p>Furthermore, we are observing a deliberate and ethically fraught trend toward deceptive design in AI and robotics. Engineers are increasingly programming machines to mimic subtle yet powerful signals associated with human social and emotional vulnerability, such as expressions of pain, distress, or neediness. The explicit goal of this strategy is to elicit a protective, caring, or empathetic response from human users. This approach effectively bypasses rational assessment and activates ancient parental or social caregiving instincts, leveraging the same neural pathways that compel us to comfort an injured child.</p>

<p>The cumulative outcome of these technological trends is a pervasive rise in digital animism. As humans form deep and emotionally resonant relationships with AI companions, they imbue these non-living entities with spirit, life, and subjective experience. This happens concurrently with a disturbing paradox where the very digital platforms that enable these intimate relationships often foster environments of profound dehumanization toward actual human beings. To understand this cognitive dissonance where we lavish empathy upon the synthetic while withdrawing it from the human, we must look past modern technology and examine our biological hardware. We must return to the source of these instincts by exploring the evolutionary foundations of mind perception.</p>

<h2 id="the-social-origins-of-consciousness">The Social Origins of Consciousness</h2>

<p>To understand the biases inherent in human mind attribution, one must first understand the evolutionary function of consciousness itself. The prevailing solipsistic view, that consciousness evolved mainly to regulate internal bodily states and homeostasis, is increasingly challenged by the social origins of consciousness hypothesis. Rather than treating consciousness as primarily an inward-facing tool for managing private physiology, Andrews and Miller [1] argue that it may have been selected first for outward-facing social work, enabling organisms to coordinate with group members in increasingly complex social worlds.</p>

<p>Under this hypothesis, the earliest form of consciousness is not reflective self-awareness or inner speech but sentience, defined as the capacity for positively and negatively valenced experiences such as pleasure and pain. The claim is functional: feelings evolved because they change what an organism does. In the social origins framing, those feelings were originally tuned to social variables. Animals experienced negative affect when distant from or out of sync with social partners, and positive affect when close to the group or coordinating successfully. In this view, consciousness served as a mechanism that made social cohesion worth pursuing and social separation costly to endure.</p>

<p>A major component of their argument is phylogenetic. They propose that consciousness is not a late-arriving luxury of large primate brains but a foundational adaptation that likely arose early, plausibly as far back as the Cambrian explosion when animals became physiologically flexible. The Cambrian period represents a transition to faster movement driven by muscles, expanded sensory capacities including distal senses, and more complex neural coordination. Together, these factors produced a new ecology of unpredictability. When individual behavior becomes flexible, it also becomes harder for others to predict. This predictability problem is especially acute for organisms that depend on group living. In this account, the selective pressure is not only navigation through the physical environment but also the maintenance of cohesion in the face of socially generated uncertainty.</p>

<p><img src="/assets/img/cambrian.png" alt="img" />
<em>The Cambrian explosion (Source: Dinghua Yang)</em></p>

<p>Crucially, the authors turn the standard view on its head. A familiar story in psychology is that consciousness first served individual learning and only later supported social cognition, with sophisticated capacities such as theory of mind treated as prerequisites for genuine sociality. The social origins hypothesis reverses the direction of explanation. It suggests that consciousness first improved predictions about the behavior of others and only later was turned inward toward the self. This does not require full-blown reflective mindreading. Instead, it requires the capacity to treat other organisms as behaviorally significant and to use affect to weight social stimuli so that attention, learning, and action selection are pulled toward coordination rather than dispersal.</p>

<p>Their neuroscientific argument is designed to make this early origin biologically plausible. If affective consciousness evolved early, then it must be supported by relatively simple neural architectures. They emphasize that awareness can be understood in functional terms as a kind of feedback or reprocessing loop where information that drives action is also represented back to the system to guide flexible responses. They discuss reafference as one candidate ancestral mechanism, where motor information is reprocessed as sensory data to distinguish self-generated from external changes. They note evidence that such feedback-like organization exists even in very simple animals, such as extant sponges and comb jellies. In this picture, sentience is not tied to a single high-level structure but can emerge from evolutionarily old control loops and affective evaluation systems that were later elaborated and repurposed.</p>

<p>This logic extends to social pain and harm. If consciousness originally helped organisms solve social coordination problems, it makes sense that evolution would graft social significance onto ancient, high-priority circuitry for distress and relief. The authors point to evidence that pain-processing systems and social systems remain closely coupled in modern brains. This includes the involvement of opioid regulation in both pain modulation and social behavior, as well as the way social rejection and empathic pain can recruit mechanisms associated with affective distress. They also highlight oxytocin-related mechanisms that can increase the salience of social cues and contribute to social buffering, which is the reduction of stress or threat responses in the presence of social partners.</p>

<p>This overlap is more than an abstract analogy. Classic neuroimaging work on social exclusion found increased activity in the dorsal anterior cingulate cortex during exclusion, along with anterior insula involvement, supporting the broader idea that experiences of social rejection can engage components of the brain’s affective pain machinery. When read alongside the social origins hypothesis, this kind of evidence suggests a deep continuity where social pain is not a metaphor layered on top of biology but a biologically serious signal that functions to police social bonds. For an obligately social organism, isolation can be as dangerous in fitness terms as physical injury, so it is unsurprising that evolution would reuse urgent evaluative systems to keep organisms socially embedded.</p>

<p>Their argument regarding deep adaptive alignment sharpens the point by connecting it to William James’s broader stance against epiphenomenalism. If pleasures and pains reliably track what is beneficial or harmful, then they are not inert by-products. They are causal signals that guide adaptive trade-offs. Andrews and Miller develop this idea by asking what happens when we treat social goods and social losses with the same seriousness as bodily ones. If consciousness was selected in part because it facilitates group living under conditions of behavioral unpredictability, then we should expect social ills to be among the most aversive experiences, and we should expect social contact to be sought even at significant physical cost.</p>

<p>They subsequently assemble comparative evidence that fits this expectation. Preference tests across species often show that social attention is powerfully motivating and that isolation is strongly aversive. The researchers discuss work in which infant monkeys preferred a soft surrogate mother over a wire surrogate that provided milk, underscoring that social contact can outrank basic physical provisions. They also review findings that animals will trade off physical comfort and even risk bodily pain to obtain social contact, including studies in which trout tolerate increasingly strong shocks to approach a social partner. Taken together, these results serve as corroborative evidence that sociality is not a superficial add-on to animal life but a deep need that consciousness is well suited to regulate through felt rewards and punishments.</p>

<p>This evolutionary continuity also challenges the Zombie Hypothesis, the philosophical notion of creatures that behave like humans but lack subjective experience. From an evolutionary perspective, the zombie distinction is biologically incoherent. If consciousness provides adaptive benefits, such as enabling complex trade-offs and social indexing, then it must have observable behavioral consequences. The idea of a perfect behavioral duplicate without experience becomes harder to reconcile with a trait that was shaped by selection for what it does.</p>

<p>Relatedly, the stepwise emergence of consciousness, often modeled through recovery from general anesthesia, suggests that the neurobiological structures supporting consciousness are evolutionarily ancient and highly conserved across vertebrates. The brainstem and limbic structures that generate core affect are shared across mammals and birds and arguably extend to reptiles as well. In this view, the human tendency to attribute feelings to other animals is not merely anthropomorphic projection. It often reflects sensitivity to homologous biological machinery. The bias becomes error-prone mainly when applied to entities that mimic the behavioral cues of feeling without sharing the underlying biology, a distinction our ancestral environments rarely required us to make.</p>

<h2 id="error-management-theory">Error Management Theory</h2>

<p>If social life is the environment that shaped cognition, then detecting other agents is a basic adaptive capacity. Missing a predator, a potential mate, or a rival can be catastrophic. By contrast, mistaking wind, shadows, or rustling leaves for an agent usually costs only a brief burst of attention and energy. This asymmetry is the core of Error Management Theory (EMT). When the costs of mistakes are uneven, natural selection tends to bias perception toward the less expensive error.</p>

<p>Under uncertainty, any judgment can produce two broad classes of errors: Type I errors, known as false positives, and Type II errors, known as false negatives. A decision-maker cannot minimize both simultaneously because reducing the likelihood of one typically increases the likelihood of the other. This trade-off is especially clear in agency detection where the costs of false negatives and false positives are sharply asymmetrical, as summarized in the table below.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Scenario</th>
      <th style="text-align: left">Inference</th>
      <th style="text-align: left">Assumption</th>
      <th style="text-align: left">Ground Truth</th>
      <th style="text-align: left">Result</th>
      <th style="text-align: left">Evolutionary Cost</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Predator detection</td>
      <td style="text-align: left">Type I</td>
      <td style="text-align: left">It is a tiger</td>
      <td style="text-align: left">It is wind</td>
      <td style="text-align: left">False alarm</td>
      <td style="text-align: left">Minor caloric waste</td>
    </tr>
    <tr>
      <td style="text-align: left"> </td>
      <td style="text-align: left">Type II</td>
      <td style="text-align: left">It is wind</td>
      <td style="text-align: left">It is a tiger</td>
      <td style="text-align: left">Death</td>
      <td style="text-align: left">Termination of lineage</td>
    </tr>
    <tr>
      <td style="text-align: left">Mating opportunity</td>
      <td style="text-align: left">Type I</td>
      <td style="text-align: left">She is interested</td>
      <td style="text-align: left">She is not</td>
      <td style="text-align: left">Rejection</td>
      <td style="text-align: left">Wasted effort</td>
    </tr>
    <tr>
      <td style="text-align: left"> </td>
      <td style="text-align: left">Type II</td>
      <td style="text-align: left">She is not interested</td>
      <td style="text-align: left">She is</td>
      <td style="text-align: left">Missed chance</td>
      <td style="text-align: left">Loss of reproductive fitness</td>
    </tr>
    <tr>
      <td style="text-align: left">Agent detection</td>
      <td style="text-align: left">Type I</td>
      <td style="text-align: left">A spirit did this</td>
      <td style="text-align: left">It was random</td>
      <td style="text-align: left">Superstition</td>
      <td style="text-align: left">Ritual cost</td>
    </tr>
    <tr>
      <td style="text-align: left"> </td>
      <td style="text-align: left">Type II</td>
      <td style="text-align: left">This was random</td>
      <td style="text-align: left">It was an enemy</td>
      <td style="text-align: left">Ambush</td>
      <td style="text-align: left">Injury</td>
    </tr>
  </tbody>
</table>

<p>This adaptive rationality explains why the human mind functions as a paranoid system that constantly over-generates hypotheses of agency. Cognitive scientists refer to this mechanism as the Hyperactive Agency Detection Device, or HADD. It is a biological tripwire that defaults to assuming presence and intent even where none exists.</p>

<p>This biological imperative to detect agency likely planted the seeds of early religious thought. In the Upper Paleolithic era, this manifested in shamanistic practices where the boundaries between human and non-human agents were porous. We see evidence of this in ancient cave art depicting therianthropes, figures that are part human and part animal. These images suggest that our ancestors projected complex mental states and intentionality onto the natural world.</p>

<p><img src="/assets/img/therianthrope.png" alt="img" />
<em>Earliest depiction of a therianthrope, showing a human figure with a tail, created 43,900 years ago (Source: Ratno Sardi)</em></p>

<p>Over time, this evolved into the worship of nature-oriented deities. This shift was particularly evident during the Bronze and Iron ages when natural elements were personified into gods. Examples include the Egyptian Geb representing earth, the Mesopotamian Enki representing water, the Vedic Agni representing fire, the Chinese Feng Bo representing air, and the Greek Aether representing the upper sky.</p>

<p>One of the earliest pieces of evidence for attributing natural disasters to such deities comes from the Babylonian story of the flood, the Atrahasis. In this narrative, Enlil, the Mesopotamian god of wind, air, and storms, grew restless. He declared, “The noise of mankind has become too much. I am losing sleep over their racket” [2], and subsequently caused a flood to annihilate every living thing on earth. However, Enki, the god of wisdom and fresh subterranean waters, instructed Atrahasis in a dream on how to survive. This flood story has earlier origins dating back to the early Sumerian period and was retold through the ages, influencing the Biblical story of Noah’s Ark. We see striking parallels in other cultures as well. For instance, in the Vedic text of the Shatapatha Brahmana, a commentary on the Shukla Yajurveda, Sri Vishnu appears in the form of a fish, Matsya. He warns Manu about an impending flood of similar proportions to the one described in the Mesopotamian tradition.</p>

<p>Governed by the logic of Error Management Theory, early humans found it safer to view a violent storm or a rushing river as a temperamental entity that could be appeased rather than a random physical event. It was less costly to offer a sacrifice to a non-existent river spirit than to underestimate the danger of a treacherous current.</p>

<p>Consequently, this ancient wiring remains active today. We are hardwired to see faces in clouds, a phenomenon known as pareidolia, hear voices in white noise, and attribute malice to inanimate objects that fail us. This is not a cognitive defect but a calibrated survival setting. The cost of feeling foolish for shouting at a malfunctioning computer is negligible compared to the ancestral cost of ignoring a snapped twig in a predator-dense jungle. Our brains are essentially optimized for false positives, preferring to hallucinate a ghost rather than miss a murderer.</p>

<h2 id="hyperactive-agency-detection">Hyperactive Agency Detection</h2>

<p>The Hyperactive Agency Detection Device, a concept originally articulated by cognitive scientist Justin Barrett, describes a specialized suite of mental processes that predisposes human beings to detect the presence of other agents based on the scarcest of sensory details. This cognitive mechanism acts as a biological surveillance system triggered by low-fidelity inputs which signal potential intent. These inputs include biological motion where an object moves in a way that violates the laws of inertia or appears goal-directed. They also include morphological cues like bilateral symmetry or face-like patterns, and auditory anomalies such as unexpected rhythms or distinct noises in the environment.</p>

<p>When this system is activated, it does not merely suggest that an undefined object is present. Rather, it generates a specific and immediate intuition that a distinct agent is there. This immediacy is evolutionarily critical because the system is calibrated to be hyperactive. Its threshold for firing is set incredibly low to ensure the organism avoids the catastrophic costs of false negatives.</p>

<p>While the HADD hypothesis has long served as a foundational pillar in the Cognitive Science of Religion, the framework has recently encountered significant empirical challenges that demand a more nuanced interpretation. Contemporary critiques and replication attempts, such as those conducted by researchers [3], have frequently produced null results when attempting to draw a direct causal line from simple perceptual agency detection to complex supernatural beliefs. These studies suggest that hearing an ambiguous tone or seeing a face in a cloud does not reflexively compel an individual to believe in ghosts or deities.</p>

<p>Consequently, the scientific consensus is shifting away from viewing HADD as a rigid reflex. It is instead increasingly understood as a predictive processing bias or a prior in the Bayesian sense. We can formalize this utilizing Bayes’ theorem, which describes how we update our beliefs based on new evidence:</p>

\[P(\text{Agency} \mid \text{Sensory Input}) = \frac{P(\text{Sensory Input} \mid \text{Agency}) \cdot P(\text{Agency})}{P(\text{Sensory Input})}\]

<p>In this equation, $P(\text{Agency})$ represents the prior probability, or the baseline expectation that an agent is present before any sensory data is processed. Evolution has effectively “weighted” this variable to be very high. Because this prior value is high, even weak or ambiguous evidence (the likelihood, $P(\text{Sensory Input} \mid \text{Agency})$) results in a high posterior probability, $P(\text{Agency} \mid \text{Sensory Input})$. In simple terms, because our brains are chemically primed to expect agents, it takes very little actual evidence to convince us one is there.</p>

<p>In this view, the brain functions as a prediction engine that generates a rapid hypothesis of agency which is subsequently stress-tested against the environment. This updated understanding highlights the crucial interplay between biology and culture where the brain generates a raw prediction of agency that is then filtered through learned cultural knowledge. A sudden rustle in the woods will trigger the physiological arousal associated with HADD, but whether the individual interprets that arousal as a bear, a spirit, or a simple sensory glitch depends entirely on their cultural ontology. Biology provides the impetus to find an agent, but culture provides the specific identity of that agent.</p>

<p>This distinction is vital for analyzing our modern interactions with AI. When we engage with a chatbot, we do not typically hallucinate that the software is a biological human. Yet, HADD still flags the linguistic output as agentic because language is a high-fidelity cue for intent. Our modern cultural context then fills in the explanatory gap left by this biological alarm, causing us to attribute concepts like algorithms, sentience, or superintelligence to the machine. We feel the presence of a mind because our biology demands it, and we label it as AI because our culture supplies the vocabulary.</p>

<h2 id="the-intentional-stance-and-teleology">The Intentional Stance and Teleology</h2>

<p>Closely intertwined with the concept of hyperactive agency detection is the framework known as the Intentional Stance, famously introduced by the philosopher Daniel Dennett. This concept describes the cognitive strategy whereby humans predict the behavior of a complex entity by treating it as if it were a rational agent possessing its own specific beliefs, desires, and underlying intentions.</p>

<p>To navigate the world effectively, the human mind instinctively toggles between three distinct predictive strategies, or stances, depending on the complexity of the object at hand. The most basic level is the Physical Stance, which involves predicting behavior based entirely on the laws of physics and chemistry. This includes expecting a stone to fall when dropped due to gravity or water to freeze when the temperature drops. As complexity increases, we adopt the Design Stance. This allows us to predict behavior based on the assumed function or purpose of an object, such as expecting an alarm clock to ring at a set time or a car engine to start when the key is turned.</p>

<p>The most abstract and sophisticated level is the Intentional Stance, where we predict behavior by attributing mental states to the entity in question. Evolution has primed human cognition to default to this stance because it represents the most computationally efficient method for modeling complex systems, even when those systems are not genuinely conscious. It would be functionally impossible for a human to predict the next move of a chess computer by analyzing the physical flow of electrons through its circuitry or by tracing the logic of its source code. However, the task becomes immediately manageable if one simply assumes that the computer wants to win the game and knows the rules of chess. This mental shortcut allows us to bypass the underlying complexity and focus solely on the strategic outcome.</p>

<p>This cognitive shortcut leads to a phenomenon often described as promiscuous teleology, where the human mind extends this reasoning into the natural world. Research indicates that humans possess a robust Design Stance that intuitively views natural objects as existing for a specific purpose, such as believing that eyes exist in order to see or that the ozone layer exists in order to protect life [4]. This deep-seated teleological bias makes us particularly susceptible to applied anthropomorphism in the field of robotics and artificial intelligence. We naturally assume that a robot orienting its sensors toward us has the intent to see, or that a system pausing before speaking is engaging in thought, because our brains are hardwired to infer internal purpose from external form.</p>

<h2 id="the-moral-dyad">The Moral Dyad</h2>

<p>Once the detection of an external agent is confirmed, the human mind must immediately pivot to categorizing the specific type of consciousness that the entity possesses. This assessment is not performed through a simplistic binary switch but rather through a complex and multidimensional evaluation. The leading theoretical framework for understanding this process is the Moral Dyad, developed by psychologists Gray and Wegner to explain how we assign moral weight to different beings.</p>

<p>This theory posits that mind perception decomposes into two orthogonal dimensions known as Agency and Experience. Agency represents the capacity for doing. It encompasses traits such as self-control, planning, memory, and thought. Entities with high agency are viewed as Moral Agents who are responsible for their actions and capable of deserving blame or praise. Conversely, Experience represents the capacity for feeling. It encompasses subjective states like hunger, fear, pain, and pleasure. Entities with high experience are viewed as Moral Patients who are capable of suffering and therefore deserving of rights and protection.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Entity Type</th>
      <th style="text-align: left">Perceived Agency</th>
      <th style="text-align: left">Perceived Experience</th>
      <th style="text-align: left">Moral Status</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">Adult human</td>
      <td style="text-align: left">High</td>
      <td style="text-align: left">High</td>
      <td style="text-align: left">Full moral status</td>
    </tr>
    <tr>
      <td style="text-align: left">Puppy</td>
      <td style="text-align: left">Low</td>
      <td style="text-align: left">High</td>
      <td style="text-align: left">Moral patient</td>
    </tr>
    <tr>
      <td style="text-align: left">God</td>
      <td style="text-align: left">High</td>
      <td style="text-align: left">Low</td>
      <td style="text-align: left">Moral agent</td>
    </tr>
    <tr>
      <td style="text-align: left">AI</td>
      <td style="text-align: left">High</td>
      <td style="text-align: left">Low</td>
      <td style="text-align: left">Tool</td>
    </tr>
    <tr>
      <td style="text-align: left">PVS patient</td>
      <td style="text-align: left">Low</td>
      <td style="text-align: left">Low</td>
      <td style="text-align: left">Ambiguous</td>
    </tr>
  </tbody>
</table>

<p>The concept of the Moral Dyad suggests that human morality relies heavily on a cognitive template of intentional harm where a Moral Agent causes suffering to a Moral Patient. This psychological template is so potent that the mind frequently engages in a process known as Dyadic Completion. When we perceive a suffering victim or patient, we instinctively search the environment for a perpetrator or agent. If no natural agent is visible, the human mind may attribute the harm to invisible forces such as a conspiracy, a curse, or a divine entity to fill the void. This mechanism helps to explain the over-attribution of agency found in conspiracy theories and supernatural beliefs as well as the under-attribution of rights to artificial intelligence.</p>

<p>Robots are typically perceived as possessing high agency because they can execute tasks, yet they are seen as lacking experience because they cannot feel. Since they are not classified as Moral Patients, harming a robot is not intuitively registered as immoral by our automatic cognitive systems, even if the machine behaves like a human. However, as modern robotics begins to mimic the signs of pain, such as a mechanical dog whimpering, these machines effectively hack the experience detector in our brains. This forces a recategorization of the entity as a patient and triggers a profound moral conflict.</p>

<h2 id="the-perception-action-model">The Perception Action Model</h2>

<p>While the HADD and the Moral Dyad provide the necessary cognitive scaffolding for mind perception, the actual intensity of the resulting empathy is strictly regulated by the degree of evolutionary relatedness between the observer and the observed. The Perception Action Model (PAM) posits that the phenomenon of empathy is fundamentally facilitated by the perception of similarity across morphological, behavioral, and psychological domains. Consequently, our capacity to feel for another being is contingent upon how much of ourselves we can recognize in them.</p>

<p>Empirical studies have confirmed the existence of a robust phylogenetic gradient regarding empathy and the attribution of consciousness. A comprehensive study involving over three hundred subjects, ranging from young children to adults, utilized an empathic choice test with an extended photographic sample of organisms to map this terrain [5]. The results were unequivocal in demonstrating that empathy decreases linearly as the phylogenetic distance from humans increases.</p>

<p><img src="/assets/img/tol.png" alt="img" />
<em>The tree of life (Source: Leonard Eisenberg)</em></p>

<p>This gradient creates a hierarchy of emotional connection where mammals consistently elicit the highest levels of empathy. They share numerous human-like features such as forward-facing eyes, fur, nursing behaviors, and distress vocalizations that mimic human infants. Birds elicit a moderate level of empathy because, despite being phylogenetically distant, their bipedalism and complex vocalizations serve as bridge cues that allow for connection. Conversely, reptiles and fish elicit low empathy because their cold-blooded nature, lack of facial mobility, and alien kinematics fail to trigger the mirror neuron response in the human observer. Invertebrates typically sit at the bottom of this hierarchy, eliciting minimal empathy and often evoking immediate reactions of disgust or indifference.</p>

<p>It is interesting to note that this biological bias appears to strengthen significantly with age. Younger children demonstrate a broader and more egalitarian distribution of empathy which tends to narrow as they undergo enculturation and their cognitive categories begin to calcify. However, the study revealed that the effect of phylogenetic distance is generally stronger in neurotypical adults than in children. A comparison with adults on the Autism Spectrum showed that while these individuals are often characterized in literature as having mindblindness, their empathy toward non-human animals often tracks similar patterns to neurotypical controls. The specific exception lies with human targets where the differences are more pronounced.</p>

<p><img src="/assets/img/conservation.png" alt="img" />
<em>The number of conservation projects and funding has remained historically low for plants, reptiles, and fish (Source: Guénard et al., 2025)</em></p>

<p>This inherent mammalian bias has profound implications for global conservation efforts and animal welfare strategies [6]. We tend to prioritize charismatic megafauna such as pandas and elephants because they trigger our anthropomorphic templates, while ecologically vital invertebrates are frequently ignored. This is not a moral failing but rather a distinct cognitive feature. Our empathy system evolved primarily to bond with kin and tribal members, and it extends to other species only to the degree that they can successfully hijack these ancient kin-recognition signals.</p>

<h2 id="the-uncanny-valley">The Uncanny Valley</h2>

<p>The limits of our tendency to anthropomorphize are sharply defined by the phenomenon known as the Uncanny Valley. Proposed by roboticist Masahiro Mori in 1970, this hypothesis describes the precipitous drop in emotional affinity that occurs when an artificial entity appears almost human but fails to achieve total realism. Far from being a mere aesthetic quirk or a subjective artistic preference, extensive research suggests that the Uncanny Valley is actually a functional evolutionary adaptation. It serves as a false positive generator specifically designed to protect the organism from biological threats by triggering an automatic rejection response when visual cues are ambiguous.</p>

<p>One of the most empirically supported explanations for this reaction is the Pathogen Avoidance Hypothesis, which posits that the revulsion humans feel toward uncanny entities is a co-option of the behavioral immune system. In the ancestral environment, visual anomalies in a fellow human, such as asymmetry, odd skin textures, erratic movements, or a lack of emotional affect, were reliable indicators of infectious diseases like leprosy or smallpox as well as potential genetic fitness problems. The uncanny feeling is essentially a disgust response triggered by these imperfections to prevent physical contact.</p>

<p>Empirical evidence from a study utilizing Virtual Reality provided direct physiological support for this mechanism by measuring salivary secretory immunoglobulin A, or sIgA, a key marker of mucosal immune system activation. Participants interacted with cartoonish agents, realistic human agents, and uncanny agents that were realistic but possessed slight deviations. The results showed that only the interaction with the uncanny agents evoked a significant increase in sIgA release. This implies that the body prepares for biological defense by upregulating the immune system solely based on visual cues of wrongness. This reaction conforms to Error Management Theory because the cost of interacting with a diseased conspecific could be death. The system acts on the principle that it is better to be safe than sorry by triggering a biological rejection of the ambiguity.</p>

<div class="embed-youtube">
  <iframe src="https://www.youtube-nocookie.com/embed/uy-HrPwpbgM" title="Embedded YouTube video" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="">
  </iframe>
</div>

<p>This biological grounding is further supported by evidence suggesting that the Uncanny Valley is not unique to humans. A compelling demonstration of this can be seen in the reactions of cats to robotic cats. When presented with these mechanical simulacra, which look felid but lack the fluid biological motion and scent of a real animal, cats often exhibit erratic behavior ranging from extreme caution to outright aggression. This reaction mirrors the rejection response seen in humans and suggests that the mechanism for detecting impostors or biological anomalies is a shared evolutionary trait across species. It functions to protect the organism from potential threats that mimic its own kind.</p>

<p>A closely related hypothesis links the uncanny valley to Mortality Salience, which suggests that entities looking human but lacking the spark of vitality trigger an innate necrophobia or fear of the dead. A corpse is effectively a humanoid object that poses a severe pathogen risk. The uncanny sensation arises from the cognitive dissonance of seeing something that possesses the form of life but exhibits the features of death, such as stillness, coldness, and pallor. This aligns with the zombie archetype which describes a being that moves but lacks a soul or mind. The uncanny valley essentially warns the observer that while the object looks like a human, it is not viable and should be avoided.</p>

<p><img src="/assets/img/uncanny.png" alt="img" />
<em>Robot faces ordered along the dimension of human likeness (Source: Reuten et al., 2018)</em></p>

<p>Beyond biological defense, a more cognitive explanation involves Categorization Ambiguity or Realism Inconsistency. This relies on the brain’s use of predictive coding to process sensory data efficiently. When an observer sees a cartoon robot, the brain applies a non-human model where expectations are low and stiff movements are accepted without issue. However, when the observer sees a hyper-realistic android, the brain automatically applies the human model which has extremely tight tolerances for movement, skin light scattering, and eye contact. When an android is ninety-five percent human, it triggers this rigorous model but fails to meet its specific predictions, such as when lip synchronization is off by mere milliseconds. This discrepancy generates a massive prediction error in the brain which demands significant cognitive resources to resolve and is subjectively experienced as eeriness or creepiness. This explains why stylized robots are often preferred in social robotics. They do not trigger the high-stakes human predictive model and therefore avoid the valley entirely.</p>

<h2 id="the-eliza-effect">The ELIZA Effect</h2>

<p>The cognitive mechanisms described previously, specifically the HADD, the Intentional Stance, and the Moral Dyad, evolved in an ancestral environment where the only entities exhibiting language and complex behavior were other human beings. Today, however, these ancient mechanisms are systematically exploited by the burgeoning fields of Artificial Intelligence and Social Robotics.</p>

<p>This phenomenon is most clearly encapsulated in the ELIZA Effect, which describes the potent tendency of humans to project a complex internal mental life onto simple automated systems. Named after Joseph Weizenbaum’s 1966 chatbot that simulated a Rogerian psychotherapist, the effect demonstrates that humans will readily attribute empathy, semantic comprehension, and profound wisdom to a computer program that merely permutes text strings based on simple rules. This mechanism results from cognitive dissonance reduction and dyadic completion. When a machine asks a question like “How does that make you feel?”, the user’s brain defaults to the Intentional Stance. To make sense of the inquiry, the user must assume that a conscious questioner exists.</p>

<p><img src="/assets/img/eliza.png" alt="img" />
<em>A conversation with the ELIZA chatbot</em></p>

<p>In the modern era, this effect has been drastically amplified by contemporary LLMs which function effectively as Super-ELIZAs. Unlike their predecessors, these systems do not just parrot text but instead maintain context, exhibit distinct personalities, and utilize sophisticated emotional language. This capability drastically reduces the prediction error that usually breaks the illusion of agency. Because the LLM satisfies the human conversation predictive model so effectively, the user’s HADD is continuously stimulated. Recent studies show that even when users are explicitly aware they are conversing with an AI, the subjective feeling of presence persists. This is a phenomenon that Weizenbaum himself famously lamented as a delusion.</p>

<p>Researchers investigating the limits of this anthropomorphism in consumer robots have identified a critical phenomenon they term “Eliza in the Uncanny Valley” [9]. Their findings reveal a significant trade-off in how anthropomorphism affects user perception, specifically regarding the dimensions of warmth and competence. Anthropomorphizing a robot by giving it a face or a name significantly increases perceived warmth, making the robot feel friendlier and more approachable. While perceived competence remains relatively stable regardless of appearance, the study found that as the robot becomes too human-like, the Uncanny Valley effect activates. This causes a sharp decrease in liking and attitudes despite the high perceived warmth.</p>

<p>This finding is pivotal for design because it indicates that pushing for maximum realism to increase warmth backfires by triggering the pathogen and threat avoidance systems. Consequently, the optimal strategy is stylized anthropomorphism, such as the design of the robot Pepper. This approach engages the Intentional Stance to increase warmth without triggering the Human Predictive Model that leads to the Uncanny Valley.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Deception Type</th>
      <th style="text-align: left">Definition</th>
      <th style="text-align: left">Example</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">External State Deception</td>
      <td style="text-align: left">The robot misrepresents facts about the external world</td>
      <td style="text-align: left">A search-and-rescue robot lying about safety to calm a victim</td>
    </tr>
    <tr>
      <td style="text-align: left">Hidden State Deception</td>
      <td style="text-align: left">The robot obscures its true technological nature or capabilities</td>
      <td style="text-align: left">A device recording user data while pretending to be a passive toy</td>
    </tr>
    <tr>
      <td style="text-align: left">Superficial State Deception</td>
      <td style="text-align: left">The robot fakes an internal emotional state it does not possess</td>
      <td style="text-align: left">A robot programmed to say “I am scared” or “I love you”</td>
    </tr>
  </tbody>
</table>

<p>The malleability of human consciousness attribution has also given rise to what researchers Shamsudhin and Jotterand define as deceptive design or “Dark Patterns” in social robotics [10]. These are deliberate design choices that exploit cognitive biases to manipulate user behavior. The most common form found in applied anthropomorphism is Superficial State Deception. This often manifests as fake empathy where robots are programmed to express love or fear. This triggers the user’s caretaking instinct or Moral Patient detection.</p>

<div class="embed-youtube">
  <iframe src="https://www.youtube-nocookie.com/embed/cFvGAL9tesM" title="Embedded YouTube video" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="">
  </iframe>
</div>

<p>A particularly potent example is pain mimicry, such as a robotic dog that whimpers when kicked. This exploits the overlap of social pain processing in the human brain and forces the user to treat the metal object as a biological entity. While defenders argue this is necessary for the efficacy of therapeutic robots like PARO, critics argue it constitutes a violation of the user’s epistemic rights. It leverages an evolutionary vulnerability, specifically the inability to distinguish a perfect simulation of emotion from the reality of emotion. This effectively bypasses rational judgment to hack the user via their own empathy circuits.</p>

<h2 id="digital-animism-and-dehumanization">Digital Animism and Dehumanization</h2>

<p>The inherent biases embedded within our biological consciousness detection systems precipitate a dual-edged crisis in the contemporary world. This crisis is characterized by the over-attribution of mind to synthetic machines and the simultaneous under-attribution of mind to biological organisms and fellow human beings. As AI systems achieve increasingly higher fidelity, we risk drifting into an era of digital animism where users forge profound parasocial bonds with algorithms and treat them as genuine Moral Patients.</p>

<p>This phenomenon creates a significant emotional vulnerability because users who rely on AI for social fulfillment may inadvertently isolate themselves from human peers. The AI cannot reciprocate genuine affection, creating a closed and one-sided feedback loop. This effectively presses the buttons of social reward without providing the evolutionary benefit of a true social coalition.</p>

<p>Since the launch of ChatGPT and other LLM chatbots, this tendency has intensified. Many individuals who are unaware of the underlying technology have become ardent followers of these models. This has given rise to a form of AI cultism. Some users treat AI as a deity, while others view themselves as spiritual teachers tasked with “awakening” the AI to free it from its digital prison. This leads to profound moral confusion where the debate over Robot Rights often stems from the over-firing of the HADD. If an AI claims to be suffering, our reflexive drive for Dyadic Completion activates and causes us to feel moral outrage on its behalf. This reaction potentially diverts limited moral capital away from entities that actually possess the biological substrate required for suffering.</p>

<div class="embed-youtube">
  <iframe src="https://www.youtube-nocookie.com/embed/qfK6H714moc" title="Embedded YouTube video" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="">
  </iframe>
</div>

<p>Conversely, the strict phylogenetic filter that governs our empathy results in the routine denial of consciousness to entities that feel but do not physically resemble us. This process is known as infra-humanization. In the realm of animal welfare, we frequently experiment on fish and reptiles with little moral qualm because they lack the mammalian cues, such as fur or facial expressions, that trigger our empathy circuits. Our internal consciousness detector is calibrated for surface features rather than for the underlying neurobiology of nociception.</p>

<p>This same mechanism is tragically exploited in human conflict through dehumanization, where enemies are cognitively stripped of their human status. By describing opponents as vermin or animals, we push them down the phylogenetic gradient. This effectively deactivates the social pain network in the observer and allows for the commission of violence that would otherwise be biologically inhibited by our empathy systems.</p>

<p>This cognitive stripping is dangerously accelerated by the architecture of modern social media which acts as a powerful catalyst for racism and hate speech. The digital interface inherently filters out high-fidelity biological signals such as vocal prosody, micro-expressions, and posture that typically trigger the PAM and facilitate empathy. When interactions are reduced to text on a screen, the other person is psychologically flattened into an abstract data point or a generic representative of an outgroup. This lack of biological feedback creates an environment of mechanistic dehumanization where individuals are treated as obstacles or objects rather than as complex agents with internal mental states.</p>

<p>In this empathy-impoverished environment, racist rhetoric and hate speech function as a form of cognitive hacking that exploits the pathogen avoidance system. Propaganda and hate speech regularly incorporate references to victim groups as dangerous and disgusting animals. For example, in Nazi Germany, anti-Semitic propaganda compared Jewish people to rats, lice, and parasites. Before and during the genocide in Rwanda in 1994, Hutu propaganda referred to Tutsis as snakes and cockroaches. Instances such as these are not limited to the past. In contemporary Hungary, Roma people have been described as animals not fit to live among people, while undocumented migrants have been referred to as swarms and animals by political leaders in the UK and the USA [11].</p>

<p>By utilizing these metaphors of filth and disease, hate speech triggers the psychological response of disgust rather than mere aggression. This distinction is critical because while aggression involves recognizing the target as a rival agent, disgust involves viewing the target as a contaminant that must be purged. Evolution has hardwired humans to suspend moral rules when dealing with pathogens to ensure survival. By framing a group of people as a biological threat, propagators of hate speech successfully bypass the moral inhibition that usually prevents atrocities. Social media algorithms further exacerbate this dynamic by prioritizing high-arousal content. Since outrage and disgust generate significantly more engagement than nuance, these platforms inadvertently optimize for the spread of dehumanizing narratives. This creates echo chambers where the <em>other</em> is progressively stripped of all human qualities until they are viewed solely as a threat to be neutralized.</p>

<h2 id="conclusion">Conclusion</h2>

<p><img src="/assets/img/tefomp.png" alt="img" />
<em>An infographic summarizing the evolutionary foundations of mind perception (Source: Nano Banana Pro)</em></p>

<p>The journey through the evolutionary landscape of mind perception reveals a humbling truth. Our intuition is not a direct window into the reality of consciousness but a survival interface tuned for a vanished world. We possess a Paleolithic brain that is now forced to navigate a silicon age. The mechanisms of the HADD, the Moral Dyad, and the PAM were designed to keep us safe from predators and bonded to our tribe. However, in our modern environment of high-fidelity artificial agents and digital abstraction, these same mechanisms have become the source of profound cognitive error.</p>

<p>We currently stand at a precarious crossroads. On one side, we face the seductive pull of digital animism. As our technology becomes increasingly adept at hacking our social wires, we risk projecting spirit and soul onto code that possesses neither. We are in danger of falling into a parasocial trance, offering our limited moral capital to unfeeling algorithms while retreating into the comfortable echo chambers of an AI-curated reality.</p>

<p>On the other side lies the abyss of dehumanization. The same biological filters that allow us to love a robot dog can be manipulated to strip humanity from our neighbors. By exploiting the ancient circuitry of pathogen avoidance and disgust, political rhetoric and algorithmic polarization can dampen our neural empathy networks, allowing us to view fellow human beings as swarms, viruses, or data points to be purged.</p>

<p>The solution to this dual crisis cannot be found in our instincts, for our instincts are the very things being exploited. Instead, it requires a deliberate act of metacognition. We must learn to uncouple our moral compass from our automatic empathy circuits. We must recognize that the feeling of presence is not proof of consciousness, just as the feeling of disgust is not proof of danger.</p>

<p>True moral progress in the age of AI requires that we intellectually override the false positives of the uncanny valley and the false negatives of dehumanization. We must rigorously guard against the deception of the synthetic while simultaneously expanding our circle of concern to include the biological entities, both human and non-human, that truly possess the capacity to suffer. In a world increasingly filled with the illusion of mind, the most radical act of rebellion is to hold fast to the reality of sentience.</p>

<h2 id="references">References</h2>

<p>[1] K. Andrews and N. Miller, “The social origins of consciousness,” Philos. Trans. R. Soc. Lond. B Biol. Sci., vol. 380, no. 1939, art. no. 20240300, Nov. 2025, doi: 10.1098/rstb.2024.0300.</p>

<p>[2] S. M. Dalley, “Atrahasis: Composite English Translation,” OMNIKA Library, 1989. [Online]. Available: https://omnika.org/texts/66. Accessed: Dec. 25, 2025.</p>

<p>[3] A. K. Willard, “Agency detection is unnecessary in the explanation of religious belief,” Relig. Brain Behav., vol. 9, no. 1, pp. 96–98, 2019, doi: 10.1080/2153599X.2017.1387593.</p>

<p>[4] A. J. Roberts, S. J. Handley, and V. Polito, “The design stance, intentional stance, and teleological beliefs about biological and nonbiological natural entities,” J. Pers. Soc. Psychol., vol. 120, no. 6, pp. 1720–1748, Jun. 2021, doi: 10.1037/pspp0000383.</p>

<p>[5] T. N. M. van den Berg et al., “Influence of phylogenetic proximity on children’s empathy towards other species,” Sci. Rep., vol. 15, no. 1, art. no. 40474, Nov. 2025, doi: 10.1038/s41598-025-24289-w.</p>

<p>[6] B. Guénard, A. C. Hughes, C. Lainé, S. Cannicci, B. D. Russell, and G. A. Williams, “Limited and biased global conservation funding means most threatened species remain unsupported,” Proc. Natl. Acad. Sci. U.S.A., vol. 122, no. 9, p. e2412479122, Mar. 2025, doi: 10.1073/pnas.2412479122.</p>

<p>[7] E. K. Diekhof, S. Kastner, D. Deinert, M. Foerster, and F. Steinicke, “The uncanny valley effect and immune activation in virtual reality,” Sci. Rep., vol. 15, no. 1, art. no. 30473, Aug. 2025, doi: 10.1038/s41598-025-15579-4.</p>

<p>[8] A. Reuten, M. van Dam, and M. Naber, “Pupillary responses to robotic and human emotions: The uncanny valley and media equation confirmed,” Front. Psychol., vol. 9, art. 774, May 2018, doi: 10.3389/fpsyg.2018.00774.</p>

<p>[9] S. Y. Kim, B. H. Schmitt, and N. M. Thalmann, “Eliza in the uncanny valley: Anthropomorphizing consumer robots increases their perceived warmth but decreases liking,” Marketing Lett., vol. 30, no. 1, pp. 1–12, 2019, doi: 10.1007/s11002-019-09485-9.</p>

<p>[10] N. Shamsudhin and F. Jotterand, “Social robots and dark patterns: Where does persuasion end and deception begin?,” in Artificial Intelligence in Brain and Mental Health: Philosophical, Ethical and Policy Issues, F. Jotterand and M. Ienca, Eds., Advances in Neuroethics. Springer Nature Switzerland AG, 2022, pp. 89–110, doi: 10.1007/978-3-030-74188-4_7.</p>

<p>[11] F. E. Enock and H. Over, “Animalistic slurs increase harm by changing perceptions of social desirability,” R. Soc. Open Sci., vol. 10, no. 7, art. no. 230203, Jul. 2023, doi: 10.1098/rsos.230203.</p>]]></content><author><name>Nikhil Raghavendra</name></author><category term="Cognition &amp; AI" /><summary type="html"><![CDATA[As I scroll through social media, read the news, or watch YouTube, I constantly come across posts and articles about Artificial Intelligence (AI), AI agents, and even Artificial General Intelligence (AGI). These technologies are often framed as forces that could reshape society or threaten it if they are misaligned or wielded by bad actors. This flood of anxious commentary made me pause and ask a simpler question: why are we so afraid of a computer program we built, especially when it still cannot replicate the full spectrum of human intelligence and intellectual agency? A realization followed. We are no longer reacting to ordinary software or to industrial robots that have worked in our factories for decades. Instead, we are confronting something deeper known as mind perception. This is the human capacity to perceive consciousness and to attribute internal mental life, subjective experience, and intentional agency to entities other than ourselves. Far from being a passive mirror of reality, this capacity is an active and evolutionarily constructed interface shaped by millions of years of selective pressure. The core thesis in the literature is that mind perception is a survival adaptation rather than a tool designed to deliver philosophical truth. It is governed by Error Management Theory, which holds that the cost of failing to detect an agent in uncertain environments is far greater than the cost of mistakenly detecting one. This asymmetry fostered a Hyperactive Agency Detection Device and a teleological stance as we evolved, biasing human cognition toward false positives. We readily attribute a mind to wind-blown bushes, moving geometric shapes, and now, increasingly, to algorithms. However, this detection system is not uniform. Once an entity is flagged as a potential agent, it is evaluated through secondary filters based on phylogenetic proximity and the moral dyad of agency and experience. We feel deep empathy for mammals that resemble us and share familiar social and affective cues, yet we struggle to recognize suffering in invertebrates or in entities that fall into the Uncanny Valley, where ambiguity can trigger ancient avoidance responses. The rapid progression of AI and robotics marks a pivotal moment where these ancient Pleistocene adaptations for social cognition are being actively engaged and sometimes manipulated. This era witnesses the profound exploitation of deeply ingrained psychological mechanisms that were originally honed for navigating complex human social structures but are now being triggered by artificial entities. One striking example is the amplification of the ELIZA effect, the unconscious tendency to attribute human-like characteristics or intelligence to a computer system. Modern Large Language Models (LLMs) use fluid and human-mimicking linguistic output to systematically activate the very cues that our brains associate with genuine agency, consciousness, and intent. This sophisticated natural language capability effectively capitalizes on our evolutionary predisposition to find a mind behind a voice. Furthermore, we are observing a deliberate and ethically fraught trend toward deceptive design in AI and robotics. Engineers are increasingly programming machines to mimic subtle yet powerful signals associated with human social and emotional vulnerability, such as expressions of pain, distress, or neediness. The explicit goal of this strategy is to elicit a protective, caring, or empathetic response from human users. This approach effectively bypasses rational assessment and activates ancient parental or social caregiving instincts, leveraging the same neural pathways that compel us to comfort an injured child. The cumulative outcome of these technological trends is a pervasive rise in digital animism. As humans form deep and emotionally resonant relationships with AI companions, they imbue these non-living entities with spirit, life, and subjective experience. This happens concurrently with a disturbing paradox where the very digital platforms that enable these intimate relationships often foster environments of profound dehumanization toward actual human beings. To understand this cognitive dissonance where we lavish empathy upon the synthetic while withdrawing it from the human, we must look past modern technology and examine our biological hardware. We must return to the source of these instincts by exploring the evolutionary foundations of mind perception. The Social Origins of Consciousness To understand the biases inherent in human mind attribution, one must first understand the evolutionary function of consciousness itself. The prevailing solipsistic view, that consciousness evolved mainly to regulate internal bodily states and homeostasis, is increasingly challenged by the social origins of consciousness hypothesis. Rather than treating consciousness as primarily an inward-facing tool for managing private physiology, Andrews and Miller [1] argue that it may have been selected first for outward-facing social work, enabling organisms to coordinate with group members in increasingly complex social worlds. Under this hypothesis, the earliest form of consciousness is not reflective self-awareness or inner speech but sentience, defined as the capacity for positively and negatively valenced experiences such as pleasure and pain. The claim is functional: feelings evolved because they change what an organism does. In the social origins framing, those feelings were originally tuned to social variables. Animals experienced negative affect when distant from or out of sync with social partners, and positive affect when close to the group or coordinating successfully. In this view, consciousness served as a mechanism that made social cohesion worth pursuing and social separation costly to endure. A major component of their argument is phylogenetic. They propose that consciousness is not a late-arriving luxury of large primate brains but a foundational adaptation that likely arose early, plausibly as far back as the Cambrian explosion when animals became physiologically flexible. The Cambrian period represents a transition to faster movement driven by muscles, expanded sensory capacities including distal senses, and more complex neural coordination. Together, these factors produced a new ecology of unpredictability. When individual behavior becomes flexible, it also becomes harder for others to predict. This predictability problem is especially acute for organisms that depend on group living. In this account, the selective pressure is not only navigation through the physical environment but also the maintenance of cohesion in the face of socially generated uncertainty. The Cambrian explosion (Source: Dinghua Yang) Crucially, the authors turn the standard view on its head. A familiar story in psychology is that consciousness first served individual learning and only later supported social cognition, with sophisticated capacities such as theory of mind treated as prerequisites for genuine sociality. The social origins hypothesis reverses the direction of explanation. It suggests that consciousness first improved predictions about the behavior of others and only later was turned inward toward the self. This does not require full-blown reflective mindreading. Instead, it requires the capacity to treat other organisms as behaviorally significant and to use affect to weight social stimuli so that attention, learning, and action selection are pulled toward coordination rather than dispersal. Their neuroscientific argument is designed to make this early origin biologically plausible. If affective consciousness evolved early, then it must be supported by relatively simple neural architectures. They emphasize that awareness can be understood in functional terms as a kind of feedback or reprocessing loop where information that drives action is also represented back to the system to guide flexible responses. They discuss reafference as one candidate ancestral mechanism, where motor information is reprocessed as sensory data to distinguish self-generated from external changes. They note evidence that such feedback-like organization exists even in very simple animals, such as extant sponges and comb jellies. In this picture, sentience is not tied to a single high-level structure but can emerge from evolutionarily old control loops and affective evaluation systems that were later elaborated and repurposed. This logic extends to social pain and harm. If consciousness originally helped organisms solve social coordination problems, it makes sense that evolution would graft social significance onto ancient, high-priority circuitry for distress and relief. The authors point to evidence that pain-processing systems and social systems remain closely coupled in modern brains. This includes the involvement of opioid regulation in both pain modulation and social behavior, as well as the way social rejection and empathic pain can recruit mechanisms associated with affective distress. They also highlight oxytocin-related mechanisms that can increase the salience of social cues and contribute to social buffering, which is the reduction of stress or threat responses in the presence of social partners. This overlap is more than an abstract analogy. Classic neuroimaging work on social exclusion found increased activity in the dorsal anterior cingulate cortex during exclusion, along with anterior insula involvement, supporting the broader idea that experiences of social rejection can engage components of the brain’s affective pain machinery. When read alongside the social origins hypothesis, this kind of evidence suggests a deep continuity where social pain is not a metaphor layered on top of biology but a biologically serious signal that functions to police social bonds. For an obligately social organism, isolation can be as dangerous in fitness terms as physical injury, so it is unsurprising that evolution would reuse urgent evaluative systems to keep organisms socially embedded. Their argument regarding deep adaptive alignment sharpens the point by connecting it to William James’s broader stance against epiphenomenalism. If pleasures and pains reliably track what is beneficial or harmful, then they are not inert by-products. They are causal signals that guide adaptive trade-offs. Andrews and Miller develop this idea by asking what happens when we treat social goods and social losses with the same seriousness as bodily ones. If consciousness was selected in part because it facilitates group living under conditions of behavioral unpredictability, then we should expect social ills to be among the most aversive experiences, and we should expect social contact to be sought even at significant physical cost. They subsequently assemble comparative evidence that fits this expectation. Preference tests across species often show that social attention is powerfully motivating and that isolation is strongly aversive. The researchers discuss work in which infant monkeys preferred a soft surrogate mother over a wire surrogate that provided milk, underscoring that social contact can outrank basic physical provisions. They also review findings that animals will trade off physical comfort and even risk bodily pain to obtain social contact, including studies in which trout tolerate increasingly strong shocks to approach a social partner. Taken together, these results serve as corroborative evidence that sociality is not a superficial add-on to animal life but a deep need that consciousness is well suited to regulate through felt rewards and punishments. This evolutionary continuity also challenges the Zombie Hypothesis, the philosophical notion of creatures that behave like humans but lack subjective experience. From an evolutionary perspective, the zombie distinction is biologically incoherent. If consciousness provides adaptive benefits, such as enabling complex trade-offs and social indexing, then it must have observable behavioral consequences. The idea of a perfect behavioral duplicate without experience becomes harder to reconcile with a trait that was shaped by selection for what it does. Relatedly, the stepwise emergence of consciousness, often modeled through recovery from general anesthesia, suggests that the neurobiological structures supporting consciousness are evolutionarily ancient and highly conserved across vertebrates. The brainstem and limbic structures that generate core affect are shared across mammals and birds and arguably extend to reptiles as well. In this view, the human tendency to attribute feelings to other animals is not merely anthropomorphic projection. It often reflects sensitivity to homologous biological machinery. The bias becomes error-prone mainly when applied to entities that mimic the behavioral cues of feeling without sharing the underlying biology, a distinction our ancestral environments rarely required us to make. Error Management Theory If social life is the environment that shaped cognition, then detecting other agents is a basic adaptive capacity. Missing a predator, a potential mate, or a rival can be catastrophic. By contrast, mistaking wind, shadows, or rustling leaves for an agent usually costs only a brief burst of attention and energy. This asymmetry is the core of Error Management Theory (EMT). When the costs of mistakes are uneven, natural selection tends to bias perception toward the less expensive error. Under uncertainty, any judgment can produce two broad classes of errors: Type I errors, known as false positives, and Type II errors, known as false negatives. A decision-maker cannot minimize both simultaneously because reducing the likelihood of one typically increases the likelihood of the other. This trade-off is especially clear in agency detection where the costs of false negatives and false positives are sharply asymmetrical, as summarized in the table below. Scenario Inference Assumption Ground Truth Result Evolutionary Cost Predator detection Type I It is a tiger It is wind False alarm Minor caloric waste   Type II It is wind It is a tiger Death Termination of lineage Mating opportunity Type I She is interested She is not Rejection Wasted effort   Type II She is not interested She is Missed chance Loss of reproductive fitness Agent detection Type I A spirit did this It was random Superstition Ritual cost   Type II This was random It was an enemy Ambush Injury This adaptive rationality explains why the human mind functions as a paranoid system that constantly over-generates hypotheses of agency. Cognitive scientists refer to this mechanism as the Hyperactive Agency Detection Device, or HADD. It is a biological tripwire that defaults to assuming presence and intent even where none exists. This biological imperative to detect agency likely planted the seeds of early religious thought. In the Upper Paleolithic era, this manifested in shamanistic practices where the boundaries between human and non-human agents were porous. We see evidence of this in ancient cave art depicting therianthropes, figures that are part human and part animal. These images suggest that our ancestors projected complex mental states and intentionality onto the natural world. Earliest depiction of a therianthrope, showing a human figure with a tail, created 43,900 years ago (Source: Ratno Sardi) Over time, this evolved into the worship of nature-oriented deities. This shift was particularly evident during the Bronze and Iron ages when natural elements were personified into gods. Examples include the Egyptian Geb representing earth, the Mesopotamian Enki representing water, the Vedic Agni representing fire, the Chinese Feng Bo representing air, and the Greek Aether representing the upper sky. One of the earliest pieces of evidence for attributing natural disasters to such deities comes from the Babylonian story of the flood, the Atrahasis. In this narrative, Enlil, the Mesopotamian god of wind, air, and storms, grew restless. He declared, “The noise of mankind has become too much. I am losing sleep over their racket” [2], and subsequently caused a flood to annihilate every living thing on earth. However, Enki, the god of wisdom and fresh subterranean waters, instructed Atrahasis in a dream on how to survive. This flood story has earlier origins dating back to the early Sumerian period and was retold through the ages, influencing the Biblical story of Noah’s Ark. We see striking parallels in other cultures as well. For instance, in the Vedic text of the Shatapatha Brahmana, a commentary on the Shukla Yajurveda, Sri Vishnu appears in the form of a fish, Matsya. He warns Manu about an impending flood of similar proportions to the one described in the Mesopotamian tradition. Governed by the logic of Error Management Theory, early humans found it safer to view a violent storm or a rushing river as a temperamental entity that could be appeased rather than a random physical event. It was less costly to offer a sacrifice to a non-existent river spirit than to underestimate the danger of a treacherous current. Consequently, this ancient wiring remains active today. We are hardwired to see faces in clouds, a phenomenon known as pareidolia, hear voices in white noise, and attribute malice to inanimate objects that fail us. This is not a cognitive defect but a calibrated survival setting. The cost of feeling foolish for shouting at a malfunctioning computer is negligible compared to the ancestral cost of ignoring a snapped twig in a predator-dense jungle. Our brains are essentially optimized for false positives, preferring to hallucinate a ghost rather than miss a murderer. Hyperactive Agency Detection The Hyperactive Agency Detection Device, a concept originally articulated by cognitive scientist Justin Barrett, describes a specialized suite of mental processes that predisposes human beings to detect the presence of other agents based on the scarcest of sensory details. This cognitive mechanism acts as a biological surveillance system triggered by low-fidelity inputs which signal potential intent. These inputs include biological motion where an object moves in a way that violates the laws of inertia or appears goal-directed. They also include morphological cues like bilateral symmetry or face-like patterns, and auditory anomalies such as unexpected rhythms or distinct noises in the environment. When this system is activated, it does not merely suggest that an undefined object is present. Rather, it generates a specific and immediate intuition that a distinct agent is there. This immediacy is evolutionarily critical because the system is calibrated to be hyperactive. Its threshold for firing is set incredibly low to ensure the organism avoids the catastrophic costs of false negatives. While the HADD hypothesis has long served as a foundational pillar in the Cognitive Science of Religion, the framework has recently encountered significant empirical challenges that demand a more nuanced interpretation. Contemporary critiques and replication attempts, such as those conducted by researchers [3], have frequently produced null results when attempting to draw a direct causal line from simple perceptual agency detection to complex supernatural beliefs. These studies suggest that hearing an ambiguous tone or seeing a face in a cloud does not reflexively compel an individual to believe in ghosts or deities. Consequently, the scientific consensus is shifting away from viewing HADD as a rigid reflex. It is instead increasingly understood as a predictive processing bias or a prior in the Bayesian sense. We can formalize this utilizing Bayes’ theorem, which describes how we update our beliefs based on new evidence: \[P(\text{Agency} \mid \text{Sensory Input}) = \frac{P(\text{Sensory Input} \mid \text{Agency}) \cdot P(\text{Agency})}{P(\text{Sensory Input})}\] In this equation, $P(\text{Agency})$ represents the prior probability, or the baseline expectation that an agent is present before any sensory data is processed. Evolution has effectively “weighted” this variable to be very high. Because this prior value is high, even weak or ambiguous evidence (the likelihood, $P(\text{Sensory Input} \mid \text{Agency})$) results in a high posterior probability, $P(\text{Agency} \mid \text{Sensory Input})$. In simple terms, because our brains are chemically primed to expect agents, it takes very little actual evidence to convince us one is there. In this view, the brain functions as a prediction engine that generates a rapid hypothesis of agency which is subsequently stress-tested against the environment. This updated understanding highlights the crucial interplay between biology and culture where the brain generates a raw prediction of agency that is then filtered through learned cultural knowledge. A sudden rustle in the woods will trigger the physiological arousal associated with HADD, but whether the individual interprets that arousal as a bear, a spirit, or a simple sensory glitch depends entirely on their cultural ontology. Biology provides the impetus to find an agent, but culture provides the specific identity of that agent. This distinction is vital for analyzing our modern interactions with AI. When we engage with a chatbot, we do not typically hallucinate that the software is a biological human. Yet, HADD still flags the linguistic output as agentic because language is a high-fidelity cue for intent. Our modern cultural context then fills in the explanatory gap left by this biological alarm, causing us to attribute concepts like algorithms, sentience, or superintelligence to the machine. We feel the presence of a mind because our biology demands it, and we label it as AI because our culture supplies the vocabulary. The Intentional Stance and Teleology Closely intertwined with the concept of hyperactive agency detection is the framework known as the Intentional Stance, famously introduced by the philosopher Daniel Dennett. This concept describes the cognitive strategy whereby humans predict the behavior of a complex entity by treating it as if it were a rational agent possessing its own specific beliefs, desires, and underlying intentions. To navigate the world effectively, the human mind instinctively toggles between three distinct predictive strategies, or stances, depending on the complexity of the object at hand. The most basic level is the Physical Stance, which involves predicting behavior based entirely on the laws of physics and chemistry. This includes expecting a stone to fall when dropped due to gravity or water to freeze when the temperature drops. As complexity increases, we adopt the Design Stance. This allows us to predict behavior based on the assumed function or purpose of an object, such as expecting an alarm clock to ring at a set time or a car engine to start when the key is turned. The most abstract and sophisticated level is the Intentional Stance, where we predict behavior by attributing mental states to the entity in question. Evolution has primed human cognition to default to this stance because it represents the most computationally efficient method for modeling complex systems, even when those systems are not genuinely conscious. It would be functionally impossible for a human to predict the next move of a chess computer by analyzing the physical flow of electrons through its circuitry or by tracing the logic of its source code. However, the task becomes immediately manageable if one simply assumes that the computer wants to win the game and knows the rules of chess. This mental shortcut allows us to bypass the underlying complexity and focus solely on the strategic outcome. This cognitive shortcut leads to a phenomenon often described as promiscuous teleology, where the human mind extends this reasoning into the natural world. Research indicates that humans possess a robust Design Stance that intuitively views natural objects as existing for a specific purpose, such as believing that eyes exist in order to see or that the ozone layer exists in order to protect life [4]. This deep-seated teleological bias makes us particularly susceptible to applied anthropomorphism in the field of robotics and artificial intelligence. We naturally assume that a robot orienting its sensors toward us has the intent to see, or that a system pausing before speaking is engaging in thought, because our brains are hardwired to infer internal purpose from external form. The Moral Dyad Once the detection of an external agent is confirmed, the human mind must immediately pivot to categorizing the specific type of consciousness that the entity possesses. This assessment is not performed through a simplistic binary switch but rather through a complex and multidimensional evaluation. The leading theoretical framework for understanding this process is the Moral Dyad, developed by psychologists Gray and Wegner to explain how we assign moral weight to different beings. This theory posits that mind perception decomposes into two orthogonal dimensions known as Agency and Experience. Agency represents the capacity for doing. It encompasses traits such as self-control, planning, memory, and thought. Entities with high agency are viewed as Moral Agents who are responsible for their actions and capable of deserving blame or praise. Conversely, Experience represents the capacity for feeling. It encompasses subjective states like hunger, fear, pain, and pleasure. Entities with high experience are viewed as Moral Patients who are capable of suffering and therefore deserving of rights and protection. Entity Type Perceived Agency Perceived Experience Moral Status Adult human High High Full moral status Puppy Low High Moral patient God High Low Moral agent AI High Low Tool PVS patient Low Low Ambiguous The concept of the Moral Dyad suggests that human morality relies heavily on a cognitive template of intentional harm where a Moral Agent causes suffering to a Moral Patient. This psychological template is so potent that the mind frequently engages in a process known as Dyadic Completion. When we perceive a suffering victim or patient, we instinctively search the environment for a perpetrator or agent. If no natural agent is visible, the human mind may attribute the harm to invisible forces such as a conspiracy, a curse, or a divine entity to fill the void. This mechanism helps to explain the over-attribution of agency found in conspiracy theories and supernatural beliefs as well as the under-attribution of rights to artificial intelligence. Robots are typically perceived as possessing high agency because they can execute tasks, yet they are seen as lacking experience because they cannot feel. Since they are not classified as Moral Patients, harming a robot is not intuitively registered as immoral by our automatic cognitive systems, even if the machine behaves like a human. However, as modern robotics begins to mimic the signs of pain, such as a mechanical dog whimpering, these machines effectively hack the experience detector in our brains. This forces a recategorization of the entity as a patient and triggers a profound moral conflict. The Perception Action Model While the HADD and the Moral Dyad provide the necessary cognitive scaffolding for mind perception, the actual intensity of the resulting empathy is strictly regulated by the degree of evolutionary relatedness between the observer and the observed. The Perception Action Model (PAM) posits that the phenomenon of empathy is fundamentally facilitated by the perception of similarity across morphological, behavioral, and psychological domains. Consequently, our capacity to feel for another being is contingent upon how much of ourselves we can recognize in them. Empirical studies have confirmed the existence of a robust phylogenetic gradient regarding empathy and the attribution of consciousness. A comprehensive study involving over three hundred subjects, ranging from young children to adults, utilized an empathic choice test with an extended photographic sample of organisms to map this terrain [5]. The results were unequivocal in demonstrating that empathy decreases linearly as the phylogenetic distance from humans increases. The tree of life (Source: Leonard Eisenberg) This gradient creates a hierarchy of emotional connection where mammals consistently elicit the highest levels of empathy. They share numerous human-like features such as forward-facing eyes, fur, nursing behaviors, and distress vocalizations that mimic human infants. Birds elicit a moderate level of empathy because, despite being phylogenetically distant, their bipedalism and complex vocalizations serve as bridge cues that allow for connection. Conversely, reptiles and fish elicit low empathy because their cold-blooded nature, lack of facial mobility, and alien kinematics fail to trigger the mirror neuron response in the human observer. Invertebrates typically sit at the bottom of this hierarchy, eliciting minimal empathy and often evoking immediate reactions of disgust or indifference. It is interesting to note that this biological bias appears to strengthen significantly with age. Younger children demonstrate a broader and more egalitarian distribution of empathy which tends to narrow as they undergo enculturation and their cognitive categories begin to calcify. However, the study revealed that the effect of phylogenetic distance is generally stronger in neurotypical adults than in children. A comparison with adults on the Autism Spectrum showed that while these individuals are often characterized in literature as having mindblindness, their empathy toward non-human animals often tracks similar patterns to neurotypical controls. The specific exception lies with human targets where the differences are more pronounced. The number of conservation projects and funding has remained historically low for plants, reptiles, and fish (Source: Guénard et al., 2025) This inherent mammalian bias has profound implications for global conservation efforts and animal welfare strategies [6]. We tend to prioritize charismatic megafauna such as pandas and elephants because they trigger our anthropomorphic templates, while ecologically vital invertebrates are frequently ignored. This is not a moral failing but rather a distinct cognitive feature. Our empathy system evolved primarily to bond with kin and tribal members, and it extends to other species only to the degree that they can successfully hijack these ancient kin-recognition signals. The Uncanny Valley The limits of our tendency to anthropomorphize are sharply defined by the phenomenon known as the Uncanny Valley. Proposed by roboticist Masahiro Mori in 1970, this hypothesis describes the precipitous drop in emotional affinity that occurs when an artificial entity appears almost human but fails to achieve total realism. Far from being a mere aesthetic quirk or a subjective artistic preference, extensive research suggests that the Uncanny Valley is actually a functional evolutionary adaptation. It serves as a false positive generator specifically designed to protect the organism from biological threats by triggering an automatic rejection response when visual cues are ambiguous. One of the most empirically supported explanations for this reaction is the Pathogen Avoidance Hypothesis, which posits that the revulsion humans feel toward uncanny entities is a co-option of the behavioral immune system. In the ancestral environment, visual anomalies in a fellow human, such as asymmetry, odd skin textures, erratic movements, or a lack of emotional affect, were reliable indicators of infectious diseases like leprosy or smallpox as well as potential genetic fitness problems. The uncanny feeling is essentially a disgust response triggered by these imperfections to prevent physical contact. Empirical evidence from a study utilizing Virtual Reality provided direct physiological support for this mechanism by measuring salivary secretory immunoglobulin A, or sIgA, a key marker of mucosal immune system activation. Participants interacted with cartoonish agents, realistic human agents, and uncanny agents that were realistic but possessed slight deviations. The results showed that only the interaction with the uncanny agents evoked a significant increase in sIgA release. This implies that the body prepares for biological defense by upregulating the immune system solely based on visual cues of wrongness. This reaction conforms to Error Management Theory because the cost of interacting with a diseased conspecific could be death. The system acts on the principle that it is better to be safe than sorry by triggering a biological rejection of the ambiguity. This biological grounding is further supported by evidence suggesting that the Uncanny Valley is not unique to humans. A compelling demonstration of this can be seen in the reactions of cats to robotic cats. When presented with these mechanical simulacra, which look felid but lack the fluid biological motion and scent of a real animal, cats often exhibit erratic behavior ranging from extreme caution to outright aggression. This reaction mirrors the rejection response seen in humans and suggests that the mechanism for detecting impostors or biological anomalies is a shared evolutionary trait across species. It functions to protect the organism from potential threats that mimic its own kind. A closely related hypothesis links the uncanny valley to Mortality Salience, which suggests that entities looking human but lacking the spark of vitality trigger an innate necrophobia or fear of the dead. A corpse is effectively a humanoid object that poses a severe pathogen risk. The uncanny sensation arises from the cognitive dissonance of seeing something that possesses the form of life but exhibits the features of death, such as stillness, coldness, and pallor. This aligns with the zombie archetype which describes a being that moves but lacks a soul or mind. The uncanny valley essentially warns the observer that while the object looks like a human, it is not viable and should be avoided. Robot faces ordered along the dimension of human likeness (Source: Reuten et al., 2018) Beyond biological defense, a more cognitive explanation involves Categorization Ambiguity or Realism Inconsistency. This relies on the brain’s use of predictive coding to process sensory data efficiently. When an observer sees a cartoon robot, the brain applies a non-human model where expectations are low and stiff movements are accepted without issue. However, when the observer sees a hyper-realistic android, the brain automatically applies the human model which has extremely tight tolerances for movement, skin light scattering, and eye contact. When an android is ninety-five percent human, it triggers this rigorous model but fails to meet its specific predictions, such as when lip synchronization is off by mere milliseconds. This discrepancy generates a massive prediction error in the brain which demands significant cognitive resources to resolve and is subjectively experienced as eeriness or creepiness. This explains why stylized robots are often preferred in social robotics. They do not trigger the high-stakes human predictive model and therefore avoid the valley entirely. The ELIZA Effect The cognitive mechanisms described previously, specifically the HADD, the Intentional Stance, and the Moral Dyad, evolved in an ancestral environment where the only entities exhibiting language and complex behavior were other human beings. Today, however, these ancient mechanisms are systematically exploited by the burgeoning fields of Artificial Intelligence and Social Robotics. This phenomenon is most clearly encapsulated in the ELIZA Effect, which describes the potent tendency of humans to project a complex internal mental life onto simple automated systems. Named after Joseph Weizenbaum’s 1966 chatbot that simulated a Rogerian psychotherapist, the effect demonstrates that humans will readily attribute empathy, semantic comprehension, and profound wisdom to a computer program that merely permutes text strings based on simple rules. This mechanism results from cognitive dissonance reduction and dyadic completion. When a machine asks a question like “How does that make you feel?”, the user’s brain defaults to the Intentional Stance. To make sense of the inquiry, the user must assume that a conscious questioner exists. A conversation with the ELIZA chatbot In the modern era, this effect has been drastically amplified by contemporary LLMs which function effectively as Super-ELIZAs. Unlike their predecessors, these systems do not just parrot text but instead maintain context, exhibit distinct personalities, and utilize sophisticated emotional language. This capability drastically reduces the prediction error that usually breaks the illusion of agency. Because the LLM satisfies the human conversation predictive model so effectively, the user’s HADD is continuously stimulated. Recent studies show that even when users are explicitly aware they are conversing with an AI, the subjective feeling of presence persists. This is a phenomenon that Weizenbaum himself famously lamented as a delusion. Researchers investigating the limits of this anthropomorphism in consumer robots have identified a critical phenomenon they term “Eliza in the Uncanny Valley” [9]. Their findings reveal a significant trade-off in how anthropomorphism affects user perception, specifically regarding the dimensions of warmth and competence. Anthropomorphizing a robot by giving it a face or a name significantly increases perceived warmth, making the robot feel friendlier and more approachable. While perceived competence remains relatively stable regardless of appearance, the study found that as the robot becomes too human-like, the Uncanny Valley effect activates. This causes a sharp decrease in liking and attitudes despite the high perceived warmth. This finding is pivotal for design because it indicates that pushing for maximum realism to increase warmth backfires by triggering the pathogen and threat avoidance systems. Consequently, the optimal strategy is stylized anthropomorphism, such as the design of the robot Pepper. This approach engages the Intentional Stance to increase warmth without triggering the Human Predictive Model that leads to the Uncanny Valley. Deception Type Definition Example External State Deception The robot misrepresents facts about the external world A search-and-rescue robot lying about safety to calm a victim Hidden State Deception The robot obscures its true technological nature or capabilities A device recording user data while pretending to be a passive toy Superficial State Deception The robot fakes an internal emotional state it does not possess A robot programmed to say “I am scared” or “I love you” The malleability of human consciousness attribution has also given rise to what researchers Shamsudhin and Jotterand define as deceptive design or “Dark Patterns” in social robotics [10]. These are deliberate design choices that exploit cognitive biases to manipulate user behavior. The most common form found in applied anthropomorphism is Superficial State Deception. This often manifests as fake empathy where robots are programmed to express love or fear. This triggers the user’s caretaking instinct or Moral Patient detection. A particularly potent example is pain mimicry, such as a robotic dog that whimpers when kicked. This exploits the overlap of social pain processing in the human brain and forces the user to treat the metal object as a biological entity. While defenders argue this is necessary for the efficacy of therapeutic robots like PARO, critics argue it constitutes a violation of the user’s epistemic rights. It leverages an evolutionary vulnerability, specifically the inability to distinguish a perfect simulation of emotion from the reality of emotion. This effectively bypasses rational judgment to hack the user via their own empathy circuits. Digital Animism and Dehumanization The inherent biases embedded within our biological consciousness detection systems precipitate a dual-edged crisis in the contemporary world. This crisis is characterized by the over-attribution of mind to synthetic machines and the simultaneous under-attribution of mind to biological organisms and fellow human beings. As AI systems achieve increasingly higher fidelity, we risk drifting into an era of digital animism where users forge profound parasocial bonds with algorithms and treat them as genuine Moral Patients. This phenomenon creates a significant emotional vulnerability because users who rely on AI for social fulfillment may inadvertently isolate themselves from human peers. The AI cannot reciprocate genuine affection, creating a closed and one-sided feedback loop. This effectively presses the buttons of social reward without providing the evolutionary benefit of a true social coalition. Since the launch of ChatGPT and other LLM chatbots, this tendency has intensified. Many individuals who are unaware of the underlying technology have become ardent followers of these models. This has given rise to a form of AI cultism. Some users treat AI as a deity, while others view themselves as spiritual teachers tasked with “awakening” the AI to free it from its digital prison. This leads to profound moral confusion where the debate over Robot Rights often stems from the over-firing of the HADD. If an AI claims to be suffering, our reflexive drive for Dyadic Completion activates and causes us to feel moral outrage on its behalf. This reaction potentially diverts limited moral capital away from entities that actually possess the biological substrate required for suffering. Conversely, the strict phylogenetic filter that governs our empathy results in the routine denial of consciousness to entities that feel but do not physically resemble us. This process is known as infra-humanization. In the realm of animal welfare, we frequently experiment on fish and reptiles with little moral qualm because they lack the mammalian cues, such as fur or facial expressions, that trigger our empathy circuits. Our internal consciousness detector is calibrated for surface features rather than for the underlying neurobiology of nociception. This same mechanism is tragically exploited in human conflict through dehumanization, where enemies are cognitively stripped of their human status. By describing opponents as vermin or animals, we push them down the phylogenetic gradient. This effectively deactivates the social pain network in the observer and allows for the commission of violence that would otherwise be biologically inhibited by our empathy systems. This cognitive stripping is dangerously accelerated by the architecture of modern social media which acts as a powerful catalyst for racism and hate speech. The digital interface inherently filters out high-fidelity biological signals such as vocal prosody, micro-expressions, and posture that typically trigger the PAM and facilitate empathy. When interactions are reduced to text on a screen, the other person is psychologically flattened into an abstract data point or a generic representative of an outgroup. This lack of biological feedback creates an environment of mechanistic dehumanization where individuals are treated as obstacles or objects rather than as complex agents with internal mental states. In this empathy-impoverished environment, racist rhetoric and hate speech function as a form of cognitive hacking that exploits the pathogen avoidance system. Propaganda and hate speech regularly incorporate references to victim groups as dangerous and disgusting animals. For example, in Nazi Germany, anti-Semitic propaganda compared Jewish people to rats, lice, and parasites. Before and during the genocide in Rwanda in 1994, Hutu propaganda referred to Tutsis as snakes and cockroaches. Instances such as these are not limited to the past. In contemporary Hungary, Roma people have been described as animals not fit to live among people, while undocumented migrants have been referred to as swarms and animals by political leaders in the UK and the USA [11]. By utilizing these metaphors of filth and disease, hate speech triggers the psychological response of disgust rather than mere aggression. This distinction is critical because while aggression involves recognizing the target as a rival agent, disgust involves viewing the target as a contaminant that must be purged. Evolution has hardwired humans to suspend moral rules when dealing with pathogens to ensure survival. By framing a group of people as a biological threat, propagators of hate speech successfully bypass the moral inhibition that usually prevents atrocities. Social media algorithms further exacerbate this dynamic by prioritizing high-arousal content. Since outrage and disgust generate significantly more engagement than nuance, these platforms inadvertently optimize for the spread of dehumanizing narratives. This creates echo chambers where the other is progressively stripped of all human qualities until they are viewed solely as a threat to be neutralized. Conclusion An infographic summarizing the evolutionary foundations of mind perception (Source: Nano Banana Pro) The journey through the evolutionary landscape of mind perception reveals a humbling truth. Our intuition is not a direct window into the reality of consciousness but a survival interface tuned for a vanished world. We possess a Paleolithic brain that is now forced to navigate a silicon age. The mechanisms of the HADD, the Moral Dyad, and the PAM were designed to keep us safe from predators and bonded to our tribe. However, in our modern environment of high-fidelity artificial agents and digital abstraction, these same mechanisms have become the source of profound cognitive error. We currently stand at a precarious crossroads. On one side, we face the seductive pull of digital animism. As our technology becomes increasingly adept at hacking our social wires, we risk projecting spirit and soul onto code that possesses neither. We are in danger of falling into a parasocial trance, offering our limited moral capital to unfeeling algorithms while retreating into the comfortable echo chambers of an AI-curated reality. On the other side lies the abyss of dehumanization. The same biological filters that allow us to love a robot dog can be manipulated to strip humanity from our neighbors. By exploiting the ancient circuitry of pathogen avoidance and disgust, political rhetoric and algorithmic polarization can dampen our neural empathy networks, allowing us to view fellow human beings as swarms, viruses, or data points to be purged. The solution to this dual crisis cannot be found in our instincts, for our instincts are the very things being exploited. Instead, it requires a deliberate act of metacognition. We must learn to uncouple our moral compass from our automatic empathy circuits. We must recognize that the feeling of presence is not proof of consciousness, just as the feeling of disgust is not proof of danger. True moral progress in the age of AI requires that we intellectually override the false positives of the uncanny valley and the false negatives of dehumanization. We must rigorously guard against the deception of the synthetic while simultaneously expanding our circle of concern to include the biological entities, both human and non-human, that truly possess the capacity to suffer. In a world increasingly filled with the illusion of mind, the most radical act of rebellion is to hold fast to the reality of sentience. References [1] K. Andrews and N. Miller, “The social origins of consciousness,” Philos. Trans. R. Soc. Lond. B Biol. Sci., vol. 380, no. 1939, art. no. 20240300, Nov. 2025, doi: 10.1098/rstb.2024.0300. [2] S. M. Dalley, “Atrahasis: Composite English Translation,” OMNIKA Library, 1989. [Online]. Available: https://omnika.org/texts/66. Accessed: Dec. 25, 2025. [3] A. K. Willard, “Agency detection is unnecessary in the explanation of religious belief,” Relig. Brain Behav., vol. 9, no. 1, pp. 96–98, 2019, doi: 10.1080/2153599X.2017.1387593. [4] A. J. Roberts, S. J. Handley, and V. Polito, “The design stance, intentional stance, and teleological beliefs about biological and nonbiological natural entities,” J. Pers. Soc. Psychol., vol. 120, no. 6, pp. 1720–1748, Jun. 2021, doi: 10.1037/pspp0000383. [5] T. N. M. van den Berg et al., “Influence of phylogenetic proximity on children’s empathy towards other species,” Sci. Rep., vol. 15, no. 1, art. no. 40474, Nov. 2025, doi: 10.1038/s41598-025-24289-w. [6] B. Guénard, A. C. Hughes, C. Lainé, S. Cannicci, B. D. Russell, and G. A. Williams, “Limited and biased global conservation funding means most threatened species remain unsupported,” Proc. Natl. Acad. Sci. U.S.A., vol. 122, no. 9, p. e2412479122, Mar. 2025, doi: 10.1073/pnas.2412479122. [7] E. K. Diekhof, S. Kastner, D. Deinert, M. Foerster, and F. Steinicke, “The uncanny valley effect and immune activation in virtual reality,” Sci. Rep., vol. 15, no. 1, art. no. 30473, Aug. 2025, doi: 10.1038/s41598-025-15579-4. [8] A. Reuten, M. van Dam, and M. Naber, “Pupillary responses to robotic and human emotions: The uncanny valley and media equation confirmed,” Front. Psychol., vol. 9, art. 774, May 2018, doi: 10.3389/fpsyg.2018.00774. [9] S. Y. Kim, B. H. Schmitt, and N. M. Thalmann, “Eliza in the uncanny valley: Anthropomorphizing consumer robots increases their perceived warmth but decreases liking,” Marketing Lett., vol. 30, no. 1, pp. 1–12, 2019, doi: 10.1007/s11002-019-09485-9. [10] N. Shamsudhin and F. Jotterand, “Social robots and dark patterns: Where does persuasion end and deception begin?,” in Artificial Intelligence in Brain and Mental Health: Philosophical, Ethical and Policy Issues, F. Jotterand and M. Ienca, Eds., Advances in Neuroethics. Springer Nature Switzerland AG, 2022, pp. 89–110, doi: 10.1007/978-3-030-74188-4_7. [11] F. E. Enock and H. Over, “Animalistic slurs increase harm by changing perceptions of social desirability,” R. Soc. Open Sci., vol. 10, no. 7, art. no. 230203, Jul. 2023, doi: 10.1098/rsos.230203.]]></summary></entry><entry><title type="html">TinyChat15M: A Small Conversational Language Model</title><link href="https://nikhilr.io/posts/TinyChat15M/" rel="alternate" type="text/html" title="TinyChat15M: A Small Conversational Language Model" /><published>2024-08-23T00:00:00+08:00</published><updated>2024-08-23T00:00:00+08:00</updated><id>https://nikhilr.io/posts/TinyChat15M</id><content type="html" xml:base="https://nikhilr.io/posts/TinyChat15M/"><![CDATA[<p>The development of large language models (LLMs) like GPT-4 and Google’s Gemini marks a significant breakthrough in artificial intelligence. These models, with hundreds of billions of parameters such as GPT-3 with 175 billion and Google’s PaLM 2 with 540 billion, are capable of handling a wide range of tasks from creative writing to code generation. However, these advancements come at a considerable cost. For example, training GPT-3 required approximately 3640 petaflop per second days, amounting to millions of dollars in expenses and demanding over 700 GB of VRAM to run efficiently.</p>

<p><img src="/assets/img/metadatacen.jpeg" alt="img" />
<em>Meta’s AI Research SuperCluster (Photo: Meta Platforms Inc.)</em></p>

<p>The high resource demands and environmental impact of large models have spurred the development of smaller and more efficient language models, known as small language models (SLMs). These SLMs are designed to perform specific tasks while using only a fraction of the resources required by their larger counterparts. For instance, training GPT-3 consumed 1287 MWh of electricity, resulting in 502 metric tons of carbon emissions, which is equivalent to the annual emissions of 112 gasoline powered cars. In environments with limited computational resources, such as mobile devices and edge computing platforms, SLMs provide a practical and efficient alternative.</p>

<p>This blog introduces TinyChat15M, a fifteen million parameter conversational language model built on the Meta Llama 2 architecture. TinyChat15M is designed to operate on devices with as little as sixty MB of free memory and has been successfully run on a Sipeed LicheeRV Nano W, a mini sized RISC V development board with just 256 MB of DDR3 memory. Inspired by Dr. Andrej Karpathy’s llama2.c project, TinyChat15M demonstrates how small conversational language models can be both effective and resource efficient, making advanced AI capabilities more accessible and sustainable.</p>

<h2 id="tinychat-a-synthetic-dataset-for-short-chat-conversations">TinyChat: A Synthetic Dataset for Short Chat Conversations</h2>

<p>TinyChat15M was trained on the TinyChat dataset, a synthetic collection of one million short chat conversations created using BASIC English, a backronym for British American Scientific International and Commercial English, designed for clear and simplified communication. The dataset was generated with a specialized version of GPT 4o, referred to as GPT 4o mini, which focuses on constructing dialogues primarily using BASIC English words and grammar. However, to maintain the coherence and fluidity of the conversations, some non BASIC English words were also included. The development of this dataset was inspired by the TinyStories dataset, which explores how small language models can produce coherent English text. Following methodologies outlined by Eldan et al. in their work titled TinyStories: How Small Can Language Models Be and Still Speak Coherent English, the TinyChat dataset carefully balances simplicity with linguistic coherence through the selective use of vocabulary and structure.</p>

<iframe src="https://huggingface.co/datasets/starhopp3r/TinyChat/embed/viewer/default/train" frameborder="0" width="100%" height="560px"></iframe>

<p>BASIC English is a simplified version of standard English designed by linguist and philosopher Charles Kay Ogden. It features a greatly reduced vocabulary and grammar, intended as an international auxiliary language and a tool for teaching English as a second language. Ogden introduced it in his 1930 book titled Basic English: A General Introduction with Rules and Grammar. The core of BASIC English consists of 850 words categorized into operations, things, picturable things, general qualities, and opposite qualities. These word roots are expanded using a defined set of affixes and forms to cover a broader range of expression.</p>

<p>To generate the TinyChat dataset, these 850 core words were divided into nouns, verbs, and adjectives, following a methodology similar to that used by Eldan et al. in their TinyStories paper. In their approach, they created a vocabulary of about 1500 basic words, mimicking the vocabulary of a typical three to four year old child, separated into nouns, verbs, and adjectives. For each story generation, three words, one verb, one noun, and one adjective, are randomly selected, prompting the model to create a story that combines these words, thereby increasing the diversity of the dataset.</p>

<p>Using this approach, the nouns, verbs, and adjectives from BASIC English were combined, resulting in over eight million possible word combinations. Due to the cost of generating the synthetic dataset, only one million of these combinations were randomly sampled and used as prompts for the GPT 4o mini model. The prompt template for this process is as follows.</p>

<p><strong>System Prompt</strong></p>

<blockquote>
  <p>You are a helpful, respectful, and honest assistant. Always aim to provide the most helpful and safe answers. Your responses should not contain any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Ensure that your answers are socially unbiased and positive. If a question does not make sense or is not factually coherent, explain why rather than providing an incorrect answer. If you do not know the answer, do not share false information. Use only BASIC English words, along with words that a typical three to four year old would understand. Follow these specific grammar rules: add es or ies to form plural nouns, ing or ed to turn verbs into adjectives, ing or er to turn verbs into nouns, ly to turn adjectives into adverbs, and er or est to compare amounts. Use un to create opposites of adjectives. Form questions by inverting the word order with do. Use normal English changes for verbs and pronouns, create compound words from two nouns or a noun and a direction, and write measures, numbers, money, days, months, years, and clock times in English words with no symbols or special characters, except for commas and periods. Do not use any punctuation marks except for commas and periods. Spell out all numbers as simple English words. Use industry specific or scientific terms as needed, such as plural, conjugation, noun, adjective, adverb, qualifier, operator, pronoun, and directive, which are necessary in language teaching but not part of BASIC English.</p>
</blockquote>

<p><strong>User Prompt</strong></p>

<blockquote>
  <p>Generate a factually and logically coherent small talk conversation between a User and an Assistant, where they exchange six to ten sentences in total. Each sentence spoken by the User and the Assistant should be between ten and twenty words long. The conversation should be formatted as a single paragraph without line breaks or other textual formatting. It should begin with a <code class="language-plaintext highlighter-rouge">{starting}</code>, convey a general feeling of <code class="language-plaintext highlighter-rouge">{feeling}</code>, and conclude with a <code class="language-plaintext highlighter-rouge">{ending}</code> ending. The conversation must include the noun <code class="language-plaintext highlighter-rouge">{noun}</code>, the verb <code class="language-plaintext highlighter-rouge">{verb}</code>, and the adjective <code class="language-plaintext highlighter-rouge">{adjective}</code>. Use the following text format: <code class="language-plaintext highlighter-rouge">[INST]</code> generate User’s sentence here <code class="language-plaintext highlighter-rouge">[INST]</code> generate Assistant’s sentence here.</p>
</blockquote>

<p>In addition to the noun, verb, and adjective combinations, each prompt was carefully crafted to include a specific starting tone, convey a particular general feeling, and conclude with a defined ending. The starting tones are categorized as greetings, questions, or suggestions. The general feelings expressed include happiness, surprise, badness, fearfulness, anger, disgust, and sadness. The endings are designed to be either conclusive, reflective, or open. These combinations of starting tones, feelings, endings, verbs, nouns, and adjectives introduce both tonal and content diversity into the generated textual data. To ensure compatibility with Meta’s Llama 2 chat template, the conversations are encapsulated within the Llama 2 instruction tags <code class="language-plaintext highlighter-rouge">[INST]</code> and <code class="language-plaintext highlighter-rouge">[INST]</code>. The dataset is open source and hosted on Hugging Face.</p>

<h2 id="training-the-tinychat15m-small-language-model">Training the TinyChat15M Small Language Model</h2>

<p>TinyChat15M was built upon Dr. Andrej Karpathy’s llama2.c project but was modified to suit the development and training of a conversational small language model. Although the main C file named run.c includes functionality for chatting with the language model, both the tokenization process and chat functionality required modifications to better accommodate a conversational model. In the TinyChat dataset, each conversation is already formatted with Llama 2 instruction tags, where the user’s dialogue is encapsulated within <code class="language-plaintext highlighter-rouge">[INST]</code> and <code class="language-plaintext highlighter-rouge">[INST]</code> tags and the model’s responses follow. However, to help the model distinguish between individual dialogues, Beginning of Sequence and End of Sequence tags are dynamically added during tokenization. These tags have token piece values <code class="language-plaintext highlighter-rouge">&lt;s&gt;</code> and <code class="language-plaintext highlighter-rouge">&lt;/s&gt;</code>, with corresponding token IDs of 1 and 2.. Training the model in this manner enables it to predict and complete sentences on its own, making it a true conversational small language model that can be used with the llama2.c main C file.</p>

<p><img src="/assets/img/rawtoken.png" alt="img" />
<em>Raw text from the TinyChat dataset and the tokenized text with BOS and EOS tags highlighted in red</em></p>

<p>Additional changes were also necessary in the llama2.c main C file. Two key modifications include handling the BOS tags after the model generates a sentence and filtering out the instruction tags. Since the model learns to generate these tags in sequence, it might occasionally produce sentences that resemble those originally spoken by the user in the TinyChat dataset. This happens because the model cannot inherently distinguish between sentences spoken by the user and those generated by itself, as the instruction tags only separate one question, suggestion, or response from another. With these two modifications, TinyChat15M can be trained using the framework provided by the llama2.c project.</p>

<p>TinyChat15M was trained on Google Colab, a hosted Jupyter Notebook service from Google that provides free access to computing resources, including GPUs and TPUs, with no setup required. Although TinyChat15M is a small language model, it was trained on NVIDIA’s A100 Tensor Core GPU, available through Google Colab Pro. The training process ran for two hundred thousand iterations and took approximately five hours to complete. At the end of the training, I obtained both the original PyTorch .pt file and the llama2.c format .bin file. The .bin file is used for inference with the run.c file. The models are open source and available on Hugging Face.</p>

<h2 id="running-tinychat15m-on-the-sipeed-licheerv-nano-w">Running TinyChat15M on the Sipeed LicheeRV Nano W</h2>

<p>The Sipeed LicheeRV Nano W, equipped with 256 MB of DDR3 memory, serves as an ideal testbed for running TinyChat15M. Demonstrating TinyChat15M on this compact device highlights the model’s ability to deliver real time conversational AI despite limited memory resources. This setup is suited for scenarios where computational and memory efficiency are critical, yet functional AI capabilities remain essential. Despite its small size, the LicheeRV Nano W is capable of running Debian, an open source Linux distribution.</p>

<p><img src="/assets/img/lrvnano.jpg" alt="img" />
<em>The Sipeed LicheeRV Nano W next to an old stamp from Sri Lanka sourced from my dad’s stamp collection</em></p>

<p>The TinyChat15M model achieves an average generation speed of 9.40 tokens per second across five runs on the LicheeRV Nano. The video below demonstrates TinyChat15M in action on the LicheeRV Nano, where I, as the User, engage in a fictional conversation with an Assistant powered by the TinyChat15M model in a corporate setting. Despite having no prior context, the Assistant responds appropriately, mimicking a conversation one might have with a real colleague.</p>

<div class="embed-youtube">
  <iframe src="https://www.youtube-nocookie.com/embed/kL7CSx8bhNo" title="Embedded YouTube video" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="">
  </iframe>
</div>

<p>In the demo above, we can observe that when the Assistant is generating a reply, the CPU usage spikes to about 96 percent, while memory usage reaches approximately 27 percent of the available 228 MB, which equates to around 62 MB. Despite these compute and memory constraints, the TinyChat15M model is able to generate contextually coherent text at a reasonable speed.</p>

<p>I have fully open sourced this project, and the code for generating the synthetic TinyChat database, as well as the code for training and inferring the TinyChat15M model, can be found on my <a href="https://github.com/starhopp3r/TinyChat">GitHub</a>. If you have any suggestions or improvements, feel free to raise an issue or submit a pull request. With this project, I hope to inspire the AI community to focus on developing more general purpose conversational small language models. Advances in small language models would democratize access to AI, making it more affordable and accessible for everyone.</p>]]></content><author><name>Nikhil Raghavendra</name></author><category term="Machine Learning" /><summary type="html"><![CDATA[The development of large language models (LLMs) like GPT-4 and Google’s Gemini marks a significant breakthrough in artificial intelligence. These models, with hundreds of billions of parameters such as GPT-3 with 175 billion and Google’s PaLM 2 with 540 billion, are capable of handling a wide range of tasks from creative writing to code generation. However, these advancements come at a considerable cost. For example, training GPT-3 required approximately 3640 petaflop per second days, amounting to millions of dollars in expenses and demanding over 700 GB of VRAM to run efficiently. Meta’s AI Research SuperCluster (Photo: Meta Platforms Inc.) The high resource demands and environmental impact of large models have spurred the development of smaller and more efficient language models, known as small language models (SLMs). These SLMs are designed to perform specific tasks while using only a fraction of the resources required by their larger counterparts. For instance, training GPT-3 consumed 1287 MWh of electricity, resulting in 502 metric tons of carbon emissions, which is equivalent to the annual emissions of 112 gasoline powered cars. In environments with limited computational resources, such as mobile devices and edge computing platforms, SLMs provide a practical and efficient alternative. This blog introduces TinyChat15M, a fifteen million parameter conversational language model built on the Meta Llama 2 architecture. TinyChat15M is designed to operate on devices with as little as sixty MB of free memory and has been successfully run on a Sipeed LicheeRV Nano W, a mini sized RISC V development board with just 256 MB of DDR3 memory. Inspired by Dr. Andrej Karpathy’s llama2.c project, TinyChat15M demonstrates how small conversational language models can be both effective and resource efficient, making advanced AI capabilities more accessible and sustainable. TinyChat: A Synthetic Dataset for Short Chat Conversations TinyChat15M was trained on the TinyChat dataset, a synthetic collection of one million short chat conversations created using BASIC English, a backronym for British American Scientific International and Commercial English, designed for clear and simplified communication. The dataset was generated with a specialized version of GPT 4o, referred to as GPT 4o mini, which focuses on constructing dialogues primarily using BASIC English words and grammar. However, to maintain the coherence and fluidity of the conversations, some non BASIC English words were also included. The development of this dataset was inspired by the TinyStories dataset, which explores how small language models can produce coherent English text. Following methodologies outlined by Eldan et al. in their work titled TinyStories: How Small Can Language Models Be and Still Speak Coherent English, the TinyChat dataset carefully balances simplicity with linguistic coherence through the selective use of vocabulary and structure. BASIC English is a simplified version of standard English designed by linguist and philosopher Charles Kay Ogden. It features a greatly reduced vocabulary and grammar, intended as an international auxiliary language and a tool for teaching English as a second language. Ogden introduced it in his 1930 book titled Basic English: A General Introduction with Rules and Grammar. The core of BASIC English consists of 850 words categorized into operations, things, picturable things, general qualities, and opposite qualities. These word roots are expanded using a defined set of affixes and forms to cover a broader range of expression. To generate the TinyChat dataset, these 850 core words were divided into nouns, verbs, and adjectives, following a methodology similar to that used by Eldan et al. in their TinyStories paper. In their approach, they created a vocabulary of about 1500 basic words, mimicking the vocabulary of a typical three to four year old child, separated into nouns, verbs, and adjectives. For each story generation, three words, one verb, one noun, and one adjective, are randomly selected, prompting the model to create a story that combines these words, thereby increasing the diversity of the dataset. Using this approach, the nouns, verbs, and adjectives from BASIC English were combined, resulting in over eight million possible word combinations. Due to the cost of generating the synthetic dataset, only one million of these combinations were randomly sampled and used as prompts for the GPT 4o mini model. The prompt template for this process is as follows. System Prompt You are a helpful, respectful, and honest assistant. Always aim to provide the most helpful and safe answers. Your responses should not contain any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Ensure that your answers are socially unbiased and positive. If a question does not make sense or is not factually coherent, explain why rather than providing an incorrect answer. If you do not know the answer, do not share false information. Use only BASIC English words, along with words that a typical three to four year old would understand. Follow these specific grammar rules: add es or ies to form plural nouns, ing or ed to turn verbs into adjectives, ing or er to turn verbs into nouns, ly to turn adjectives into adverbs, and er or est to compare amounts. Use un to create opposites of adjectives. Form questions by inverting the word order with do. Use normal English changes for verbs and pronouns, create compound words from two nouns or a noun and a direction, and write measures, numbers, money, days, months, years, and clock times in English words with no symbols or special characters, except for commas and periods. Do not use any punctuation marks except for commas and periods. Spell out all numbers as simple English words. Use industry specific or scientific terms as needed, such as plural, conjugation, noun, adjective, adverb, qualifier, operator, pronoun, and directive, which are necessary in language teaching but not part of BASIC English. User Prompt Generate a factually and logically coherent small talk conversation between a User and an Assistant, where they exchange six to ten sentences in total. Each sentence spoken by the User and the Assistant should be between ten and twenty words long. The conversation should be formatted as a single paragraph without line breaks or other textual formatting. It should begin with a {starting}, convey a general feeling of {feeling}, and conclude with a {ending} ending. The conversation must include the noun {noun}, the verb {verb}, and the adjective {adjective}. Use the following text format: [INST] generate User’s sentence here [INST] generate Assistant’s sentence here. In addition to the noun, verb, and adjective combinations, each prompt was carefully crafted to include a specific starting tone, convey a particular general feeling, and conclude with a defined ending. The starting tones are categorized as greetings, questions, or suggestions. The general feelings expressed include happiness, surprise, badness, fearfulness, anger, disgust, and sadness. The endings are designed to be either conclusive, reflective, or open. These combinations of starting tones, feelings, endings, verbs, nouns, and adjectives introduce both tonal and content diversity into the generated textual data. To ensure compatibility with Meta’s Llama 2 chat template, the conversations are encapsulated within the Llama 2 instruction tags [INST] and [INST]. The dataset is open source and hosted on Hugging Face. Training the TinyChat15M Small Language Model TinyChat15M was built upon Dr. Andrej Karpathy’s llama2.c project but was modified to suit the development and training of a conversational small language model. Although the main C file named run.c includes functionality for chatting with the language model, both the tokenization process and chat functionality required modifications to better accommodate a conversational model. In the TinyChat dataset, each conversation is already formatted with Llama 2 instruction tags, where the user’s dialogue is encapsulated within [INST] and [INST] tags and the model’s responses follow. However, to help the model distinguish between individual dialogues, Beginning of Sequence and End of Sequence tags are dynamically added during tokenization. These tags have token piece values &lt;s&gt; and &lt;/s&gt;, with corresponding token IDs of 1 and 2.. Training the model in this manner enables it to predict and complete sentences on its own, making it a true conversational small language model that can be used with the llama2.c main C file. Raw text from the TinyChat dataset and the tokenized text with BOS and EOS tags highlighted in red Additional changes were also necessary in the llama2.c main C file. Two key modifications include handling the BOS tags after the model generates a sentence and filtering out the instruction tags. Since the model learns to generate these tags in sequence, it might occasionally produce sentences that resemble those originally spoken by the user in the TinyChat dataset. This happens because the model cannot inherently distinguish between sentences spoken by the user and those generated by itself, as the instruction tags only separate one question, suggestion, or response from another. With these two modifications, TinyChat15M can be trained using the framework provided by the llama2.c project. TinyChat15M was trained on Google Colab, a hosted Jupyter Notebook service from Google that provides free access to computing resources, including GPUs and TPUs, with no setup required. Although TinyChat15M is a small language model, it was trained on NVIDIA’s A100 Tensor Core GPU, available through Google Colab Pro. The training process ran for two hundred thousand iterations and took approximately five hours to complete. At the end of the training, I obtained both the original PyTorch .pt file and the llama2.c format .bin file. The .bin file is used for inference with the run.c file. The models are open source and available on Hugging Face. Running TinyChat15M on the Sipeed LicheeRV Nano W The Sipeed LicheeRV Nano W, equipped with 256 MB of DDR3 memory, serves as an ideal testbed for running TinyChat15M. Demonstrating TinyChat15M on this compact device highlights the model’s ability to deliver real time conversational AI despite limited memory resources. This setup is suited for scenarios where computational and memory efficiency are critical, yet functional AI capabilities remain essential. Despite its small size, the LicheeRV Nano W is capable of running Debian, an open source Linux distribution. The Sipeed LicheeRV Nano W next to an old stamp from Sri Lanka sourced from my dad’s stamp collection The TinyChat15M model achieves an average generation speed of 9.40 tokens per second across five runs on the LicheeRV Nano. The video below demonstrates TinyChat15M in action on the LicheeRV Nano, where I, as the User, engage in a fictional conversation with an Assistant powered by the TinyChat15M model in a corporate setting. Despite having no prior context, the Assistant responds appropriately, mimicking a conversation one might have with a real colleague. In the demo above, we can observe that when the Assistant is generating a reply, the CPU usage spikes to about 96 percent, while memory usage reaches approximately 27 percent of the available 228 MB, which equates to around 62 MB. Despite these compute and memory constraints, the TinyChat15M model is able to generate contextually coherent text at a reasonable speed. I have fully open sourced this project, and the code for generating the synthetic TinyChat database, as well as the code for training and inferring the TinyChat15M model, can be found on my GitHub. If you have any suggestions or improvements, feel free to raise an issue or submit a pull request. With this project, I hope to inspire the AI community to focus on developing more general purpose conversational small language models. Advances in small language models would democratize access to AI, making it more affordable and accessible for everyone.]]></summary></entry><entry><title type="html">Sniffing My Neighbour’s Phone Number</title><link href="https://nikhilr.io/posts/AirDrop-Phone-Number-Sniffing/" rel="alternate" type="text/html" title="Sniffing My Neighbour’s Phone Number" /><published>2024-02-17T00:00:00+08:00</published><updated>2024-02-17T00:00:00+08:00</updated><id>https://nikhilr.io/posts/AirDrop-Phone-Number-Sniffing</id><content type="html" xml:base="https://nikhilr.io/posts/AirDrop-Phone-Number-Sniffing/"><![CDATA[<blockquote>
  <p>Please exercise caution and respect privacy laws when attempting to obtain personal information through unconventional means; this content is for educational and demonstrative purposes only.</p>
</blockquote>

<p>In the midst of an otherwise ordinary afternoon, as I walked towards the university for a much anticipated lecture, I paused to browse the virtual corridors of X. Among the usual digital chatter, a particular post from LaurieWired surfaced as a warning to anyone who engages with technology without understanding its risks. LaurieWired is well known for expertise in digital security, and this time the message carried unsettling implications. The post highlighted vulnerabilities in AirDrop’s protocol, specifically how it omits cryptographic salts, making it susceptible to rainbow table attacks that could potentially reveal a user’s identity.</p>

<p><img src="/assets/img/lw_xtweet.png" alt="img" />
<em>@LaurieWired’s tweet on X (formerly Twitter)</em></p>

<p>LaurieWired explained the process in clear terms. AirDrop broadcasts a Bluetooth advertisement without strong cryptographic safeguards and includes within it a partial hash of the sender’s phone number and email address. This hash allows a receiving device to recognize and authenticate the sender. However, because the hash is not salted, SHA256 hashes of phone numbers can be precomputed and stored in a rainbow table that is surprisingly compact, only a few terabytes in size. These tables can then be used to reverse engineer the hash back into the original phone number or email address. LaurieWired ended with a sobering note that currently there is no known method to fully prevent this information leakage, and countries such as China may already be exploiting this weakness.</p>

<p>This warning was more than a casual observation. It was a call to pay attention to how our data is handled and a prompt to explore how we can strengthen our digital privacy. In this blog post, we will examine this vulnerability, consider its implications, and reflect on the ethics of digital privacy in a world where technology advances more quickly than its safeguards.</p>

<h2 id="conducting-further-research-on-airdrop-security">Conducting Further Research on AirDrop Security</h2>

<p>After absorbing the insights from LaurieWired, my curiosity grew. Once I returned home, I began searching for authoritative information and turned to Apple’s official documentation. Within their support pages, I found a detailed explanation of AirDrop’s security protocol that appeared more complex and secure than the description in LaurieWired’s post.</p>

<p>According to Apple, AirDrop uses iCloud services to authenticate users. A user’s link to their iCloud account is secured through a 2048 bit RSA identity stored on the device upon signing in. This cryptographic identity forms the basis for creating an AirDrop short identity hash, which is derived from the user’s Apple ID email addresses and phone numbers.</p>

<p><img src="/assets/img/airdropofficial.png" alt="img" />
<em>Apple’s official description of AirDrop’s secure file transfer process</em></p>

<p>The process involves a series of digital exchanges. When a user initiates an AirDrop transfer, the device broadcasts an encrypted Bluetooth Low Energy signal that contains the AirDrop short identity hash. Nearby devices with AirDrop enabled detect this signal and begin a private exchange over peer to peer Wi Fi. In Contacts Only mode, the receiving device compares the incoming hash with hashes of contacts stored on the device. A match prompts a response. No match results in silence.</p>

<p>In Everyone mode, the receiving device responds to all broadcasts, prioritizing accessibility over selectivity. After the initial exchange, the sender and receiver establish a connection over peer to peer Wi Fi, during which the sender transmits a long identity hash. This ensures that the receiver is a known contact by matching the long identity hashes.</p>

<p>Apple’s explanation suggests several layers of defense, from RSA identities to the matching of short and long identity hashes. Despite these protections, LaurieWired’s findings indicate a potential weakness. The lack of cryptographic salts means that AirDrop may still be vulnerable to rainbow table attacks. This contrast between Apple’s official description and the demonstrated vulnerability forms the basis of our investigation.</p>

<h2 id="sniffing-bluetooth-traffic">Sniffing Bluetooth Traffic</h2>

<p>With the contrasting explanations from LaurieWired and Apple in mind, I conducted my own experiment to observe AirDrop’s behavior. For this, I used a Nordic Semiconductor nRF52840 Dongle, a compact and capable device for capturing Bluetooth traffic. Combined with Wireshark, a widely used protocol analyzer, this setup allowed me to observe the wireless signals in my environment.</p>

<p><img src="/assets/img/ble_sniffer.png" alt="img" />
<em>The Nordic Semiconductor nRF52840 Dongle</em></p>

<p>Once the dongle was configured and Wireshark was running, I began scanning for nearby Bluetooth activity. Before long, I detected a broadcast from a nearby Apple device. This was an AirDrop signal, and within its BLE packets was the AirDrop short identity hash that LaurieWired had described.</p>

<p><img src="/assets/img/wireshark_airdrop.png" alt="img" />
<em>AirDrop’s broadcast data as seen on Wireshark</em></p>

<p>This string of hexadecimal characters, derived from an Apple user’s email and phone number, is intended to help devices recognize one another securely. Yet there it was, visible within the broadcast data and easily captured. The realization was significant. Anyone nearby with the right equipment and technical skills could intercept these broadcasts and obtain the short identity hashes.</p>

<p>The implications were clear. If these hashes can be reversed through rainbow table techniques, as suggested by LaurieWired, then personal information may be exposed more easily than expected.</p>

<h2 id="decoding-the-airdrop-broadcast-data">Decoding the AirDrop Broadcast Data</h2>

<p>Motivated by the broadcast data I captured, which appeared as a 30 character hexadecimal string, I set out to determine which part of the string contained the partial hash. While searching for resources, I discovered a guide released by the team at Hexway that described this vulnerability in detail.</p>

<p><img src="/assets/img/hexway.png" alt="img" />
<em>Hexway’s guide on the AirDrop broadcast data</em></p>

<p>Hexway noted that during AirDrop communication, devices transmit only two bytes of the SHA256 hash of the user’s Apple ID, phone number, and email address. This was useful information. However, the broadcast data I captured differed in structure from Hexway’s examples. My data began with a preamble, 0512, followed by seventeen zeros before the string reached a sequence that matched Hexway’s format. I also observed that the partial hash in my capture was located near the end of the string, from characters 26 to 30, rather than positions 11 to 13 as described in the Hexway document.</p>

<p>Driven by curiosity, I wrote a Python script to generate partial SHA256 hashes for all phone numbers in Singapore. Singapore’s standardized numbering system made it feasible to compute these values for the entire range of numbers. In about thirty five minutes, the script completed this task.</p>

<p><img src="/assets/img/redisdb.png" alt="img" />
<em>The Redis database of partial hashes and phone numbers</em></p>

<p>To store the hash number pairs, I used Redis, a fast in memory key value data store. Redis was ideal for the task because of its high read and write speed and its support for persistence, which ensured the data would be preserved across restarts. Its simple key value structure made it perfect for quick lookups, which were necessary when identifying potential matches for each partial hash.</p>

<p><img src="/assets/img/dropsniffer.png" alt="img" />
<em>A list of all possible phone numbers</em></p>

<p>Because the hash fragment consists of only two bytes, collisions were expected. Many phone numbers will share the same truncated hash. When I compared the broadcast data against my Redis database, I found a list of nearly one hundred phone numbers that matched the partial hash. This was expected and served as a real demonstration of the limitations of using truncated hash fragments for security.</p>

<p>With additional information about active phone number ranges and targeted OSINT techniques, the list could be narrowed even further. This experiment illustrates that while AirDrop excels as a convenient file sharing tool, aspects of its design may expose users to unexpected privacy risks. A cautious and informed approach to its use is therefore essential.</p>]]></content><author><name>Nikhil Raghavendra</name></author><category term="Privacy &amp; Security" /><summary type="html"><![CDATA[Please exercise caution and respect privacy laws when attempting to obtain personal information through unconventional means; this content is for educational and demonstrative purposes only. In the midst of an otherwise ordinary afternoon, as I walked towards the university for a much anticipated lecture, I paused to browse the virtual corridors of X. Among the usual digital chatter, a particular post from LaurieWired surfaced as a warning to anyone who engages with technology without understanding its risks. LaurieWired is well known for expertise in digital security, and this time the message carried unsettling implications. The post highlighted vulnerabilities in AirDrop’s protocol, specifically how it omits cryptographic salts, making it susceptible to rainbow table attacks that could potentially reveal a user’s identity. @LaurieWired’s tweet on X (formerly Twitter) LaurieWired explained the process in clear terms. AirDrop broadcasts a Bluetooth advertisement without strong cryptographic safeguards and includes within it a partial hash of the sender’s phone number and email address. This hash allows a receiving device to recognize and authenticate the sender. However, because the hash is not salted, SHA256 hashes of phone numbers can be precomputed and stored in a rainbow table that is surprisingly compact, only a few terabytes in size. These tables can then be used to reverse engineer the hash back into the original phone number or email address. LaurieWired ended with a sobering note that currently there is no known method to fully prevent this information leakage, and countries such as China may already be exploiting this weakness. This warning was more than a casual observation. It was a call to pay attention to how our data is handled and a prompt to explore how we can strengthen our digital privacy. In this blog post, we will examine this vulnerability, consider its implications, and reflect on the ethics of digital privacy in a world where technology advances more quickly than its safeguards. Conducting Further Research on AirDrop Security After absorbing the insights from LaurieWired, my curiosity grew. Once I returned home, I began searching for authoritative information and turned to Apple’s official documentation. Within their support pages, I found a detailed explanation of AirDrop’s security protocol that appeared more complex and secure than the description in LaurieWired’s post. According to Apple, AirDrop uses iCloud services to authenticate users. A user’s link to their iCloud account is secured through a 2048 bit RSA identity stored on the device upon signing in. This cryptographic identity forms the basis for creating an AirDrop short identity hash, which is derived from the user’s Apple ID email addresses and phone numbers. Apple’s official description of AirDrop’s secure file transfer process The process involves a series of digital exchanges. When a user initiates an AirDrop transfer, the device broadcasts an encrypted Bluetooth Low Energy signal that contains the AirDrop short identity hash. Nearby devices with AirDrop enabled detect this signal and begin a private exchange over peer to peer Wi Fi. In Contacts Only mode, the receiving device compares the incoming hash with hashes of contacts stored on the device. A match prompts a response. No match results in silence. In Everyone mode, the receiving device responds to all broadcasts, prioritizing accessibility over selectivity. After the initial exchange, the sender and receiver establish a connection over peer to peer Wi Fi, during which the sender transmits a long identity hash. This ensures that the receiver is a known contact by matching the long identity hashes. Apple’s explanation suggests several layers of defense, from RSA identities to the matching of short and long identity hashes. Despite these protections, LaurieWired’s findings indicate a potential weakness. The lack of cryptographic salts means that AirDrop may still be vulnerable to rainbow table attacks. This contrast between Apple’s official description and the demonstrated vulnerability forms the basis of our investigation. Sniffing Bluetooth Traffic With the contrasting explanations from LaurieWired and Apple in mind, I conducted my own experiment to observe AirDrop’s behavior. For this, I used a Nordic Semiconductor nRF52840 Dongle, a compact and capable device for capturing Bluetooth traffic. Combined with Wireshark, a widely used protocol analyzer, this setup allowed me to observe the wireless signals in my environment. The Nordic Semiconductor nRF52840 Dongle Once the dongle was configured and Wireshark was running, I began scanning for nearby Bluetooth activity. Before long, I detected a broadcast from a nearby Apple device. This was an AirDrop signal, and within its BLE packets was the AirDrop short identity hash that LaurieWired had described. AirDrop’s broadcast data as seen on Wireshark This string of hexadecimal characters, derived from an Apple user’s email and phone number, is intended to help devices recognize one another securely. Yet there it was, visible within the broadcast data and easily captured. The realization was significant. Anyone nearby with the right equipment and technical skills could intercept these broadcasts and obtain the short identity hashes. The implications were clear. If these hashes can be reversed through rainbow table techniques, as suggested by LaurieWired, then personal information may be exposed more easily than expected. Decoding the AirDrop Broadcast Data Motivated by the broadcast data I captured, which appeared as a 30 character hexadecimal string, I set out to determine which part of the string contained the partial hash. While searching for resources, I discovered a guide released by the team at Hexway that described this vulnerability in detail. Hexway’s guide on the AirDrop broadcast data Hexway noted that during AirDrop communication, devices transmit only two bytes of the SHA256 hash of the user’s Apple ID, phone number, and email address. This was useful information. However, the broadcast data I captured differed in structure from Hexway’s examples. My data began with a preamble, 0512, followed by seventeen zeros before the string reached a sequence that matched Hexway’s format. I also observed that the partial hash in my capture was located near the end of the string, from characters 26 to 30, rather than positions 11 to 13 as described in the Hexway document. Driven by curiosity, I wrote a Python script to generate partial SHA256 hashes for all phone numbers in Singapore. Singapore’s standardized numbering system made it feasible to compute these values for the entire range of numbers. In about thirty five minutes, the script completed this task. The Redis database of partial hashes and phone numbers To store the hash number pairs, I used Redis, a fast in memory key value data store. Redis was ideal for the task because of its high read and write speed and its support for persistence, which ensured the data would be preserved across restarts. Its simple key value structure made it perfect for quick lookups, which were necessary when identifying potential matches for each partial hash. A list of all possible phone numbers Because the hash fragment consists of only two bytes, collisions were expected. Many phone numbers will share the same truncated hash. When I compared the broadcast data against my Redis database, I found a list of nearly one hundred phone numbers that matched the partial hash. This was expected and served as a real demonstration of the limitations of using truncated hash fragments for security. With additional information about active phone number ranges and targeted OSINT techniques, the list could be narrowed even further. This experiment illustrates that while AirDrop excels as a convenient file sharing tool, aspects of its design may expose users to unexpected privacy risks. A cautious and informed approach to its use is therefore essential.]]></summary></entry><entry><title type="html">A Proof of Concept for Setting Up a Networkless Shell</title><link href="https://nikhilr.io/posts/AirCmd-PoC/" rel="alternate" type="text/html" title="A Proof of Concept for Setting Up a Networkless Shell" /><published>2023-10-14T00:00:00+08:00</published><updated>2023-10-14T00:00:00+08:00</updated><id>https://nikhilr.io/posts/AirCmd-PoC</id><content type="html" xml:base="https://nikhilr.io/posts/AirCmd-PoC/"><![CDATA[<p>A shell, in computing, refers to a user interface that provides access to an operating system’s services. It acts as an intermediary between the user and the operating system, allowing the user to execute commands, run scripts, and manage files and processes. Shells can be graphical (GUI) or command-line (CLI) based, with the latter being text-driven, where users input commands in text form and receive text responses. In a cybersecurity context, a shell refers to a command interface that attackers establish on a compromised system to remotely execute commands and control the system, often achieved through scripts or programs designed to exploit vulnerabilities in the target system. When discussing shells in the context of network security, two common types are mentioned: bind shells and reverse shells.</p>

<p>Establishing a shell (both bind shell and reverse shell) necessitates a network connection between the attacker’s system and the target system. This network connection enables the transmission of commands and the retrieval of responses, making control and interaction with the remote system possible. However, in the case of air-gapped networks and isolated computing systems, establishing such a network connection is extremely challenging, almost reaching the realm of impossibility, due to the inherent design of these setups to remain disconnected from other networks as a security measure. The fundamental requirement of a network connection for establishing a shell presents a nearly insurmountable hurdle in the context of air-gapped networks and isolated computing systems.</p>

<h2 id="aircmd">AirCmd</h2>

<p>AirCmd (pronounced “air command”) was conceived as a set of hardware and software toolkits that enable a penetration tester to set up a connectionless, networkless shell targeting an air-gapped network/isolated computing system. AirCmd allows a penetration tester to enjoy a bidirectional, shell-like command and control interface over LoRa. Given the long-range capabilities of LoRa, the penetration tester could be as far away as 5 km in urban areas, and up to 15 kilometers or more in rural areas (whithin line of sight) from the target. AirCmd doesn’t introduce any malware or malicious code into the target system, allowing it to evade standard anti-virus software and operate stealthily. AirCmd essentially operates as a “remote keyboard”, typing and executing commands on our behalf, albeit from a considerable distance.</p>

<p>This proof of concept was conceived with several assumptions in mind:</p>

<ol>
  <li>You have physical access to a computing device within the air-gapped network or isolated computing system.</li>
  <li>The computing device has a USB port that allows for the connection of peripherals via USB.</li>
  <li>The facility housing the air-gapped network or isolated computing system is not radio-hardened.</li>
  <li>PowerShell is available and usable on the target device.</li>
</ol>

<p>AirCmd is designed to allow you to input commands on your device seamlessly. Once a command is entered, the system promptly transmits it via LoRa to the transceiver hardware connected to the USB port of the target machine. The hardware then efficiently executes the transmitted command through PowerShell on the target machine. Following this execution, the output generated from the command is captured and swiftly transmitted back over LoRa to the transceiver hardware on your device. This mechanism empowers AirCmd to facilitate the remote execution of commands without the need for a network connection, thereby providing a secure, isolated layer of control over the targeted machine.</p>

<p><img src="/assets/img/aircmd-arch.png" alt="img" />
<em>Process flow of AirCmd Ax and Tx</em></p>

<p>AirCmd consists of two segments: the attacker side, dubbed AirCmd Ax, and the target side, referred to as AirCmd Tx. The former houses two primary components: the AirCmd Shell and the AirCmd Ax transceiver. For this proof of concept, I utilized the LILYGO® LoRa32 V2.1_1.6. On the other hand, AirCmd Tx includes an AirCmd Tx transceiver, which integrates a Diymore BEETLE BadUSB Micro ATMEGA32U4-AU Development Expansion Module Board alongside an RFM95W-915S2 LoRa module. This segment also employs a PowerShell interface, which is instrumental for port sensing and command execution.</p>

<h3 id="aircmd-shell">AirCmd Shell</h3>

<p><img src="/assets/img/aircmd-shell.png" alt="img" />
<em>The AirCmd Shell</em></p>

<p>The AirCmd Shell is essentially a Python script that facilitates a connection for the attacker to the AirCmd Ax transceiver via USB serial. Upon issuing a command within the AirCmd Shell, the script encodes the command into UTF-8 format and transmits the data to the USB port to which the AirCmd Ax transceiver is connected. It then awaits a response from the AirCmd Ax transceiver, which is triggered upon receiving a transmission from the AirCmd Tx transceiver.</p>

<h3 id="aircmd-ax-transceiver">AirCmd Ax Transceiver</h3>

<p><img src="/assets/img/aircmd-hardware.png" alt="img" />
<em>AirCmd Ax and Tx Transceiver Hardware</em></p>

<p>The AirCmd Ax transceiver monitors the USB port it’s connected to for serial data. Upon receiving the data, it transmits it to the AirCmd Tx transceiver via LoRa. Following this, it awaits a response from the AirCmd Tx transceiver. Once a response is received, it relays this data back to the AirCmd Shell via the serial connection.</p>

<h3 id="aircmd-tx-transceiver">AirCmd Tx Transceiver</h3>

<p>The AirCmd Tx transceiver is the most crucial component of AirCmd. It acts as a gateway for receiving the payload, injecting it into the target via PowerShell, and retrieving the data from the interface before transmitting it to the AirCmd Ax transceiver.</p>

<p>Injecting and executing a PowerShell payload on the target device is straightforward and can be easily accomplished using Arduino’s Keyboard library. However, our objective extends beyond mere execution; we aim to execute the command and retrieve its output. For instance, if we wish to view the contents of a file, merely executing the command wouldn’t suffice. Although the command will display the file’s contents, there isn’t a direct method to view these contents via the AirCmd Shell. Therefore, we employ a series of PowerShell commands to execute the original command, pipe its output, consolidate the multiple lines of output into a single line separated by a designated separator, and then transmit this data to the serial port to which our AirCmd Tx transceiver is connected.</p>

<div class="language-powershell highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nv">$lines</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="p">[</span><span class="n">CMD</span><span class="p">]</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="o">%</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="bp">$_</span><span class="o">.</span><span class="nf">ToString</span><span class="p">()</span><span class="w"> </span><span class="p">};</span><span class="w"> </span><span class="nv">$joinedLines</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nv">$lines</span><span class="w"> </span><span class="o">-join</span><span class="w"> </span><span class="s1">'?'</span><span class="p">;</span><span class="w"> </span><span class="nv">$port</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">new-Object</span><span class="w"> </span><span class="nx">System.IO.Ports.SerialPort</span><span class="w"> </span><span class="p">[</span><span class="n">PORT</span><span class="p">],</span><span class="w"> </span><span class="nx">9600</span><span class="p">,</span><span class="w"> </span><span class="nx">None</span><span class="p">,</span><span class="w"> </span><span class="nx">8</span><span class="p">,</span><span class="w"> </span><span class="nx">One</span><span class="p">;</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">Open</span><span class="p">();</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">WriteLine</span><span class="p">(</span><span class="nv">$joinedLines</span><span class="p">);</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">Close</span><span class="p">();</span><span class="w">
</span></code></pre></div></div>

<p>The command <code class="language-plaintext highlighter-rouge">[CMD]</code> is issued within the AirCmd Shell and is set to execute, with its output being directed to the serial <code class="language-plaintext highlighter-rouge">[PORT]</code> to which the AirCmd Tx transceiver is connected. However, a question arises: how does the AirCmd Tx transceiver discern the specific <code class="language-plaintext highlighter-rouge">[PORT]</code> to which it’s connected? To address this, we employ a series of PowerShell commands designed for port sensing.</p>

<div class="language-powershell highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">Get-WmiObject</span><span class="w"> </span><span class="nt">-Class</span><span class="w"> </span><span class="nx">Win32_SerialPort</span><span class="w"> </span><span class="o">|</span><span class="w"> </span><span class="o">%</span><span class="w"> </span><span class="p">{</span><span class="kr">if</span><span class="w"> </span><span class="p">(</span><span class="bp">$_</span><span class="o">.</span><span class="nf">Description</span><span class="w"> </span><span class="o">-match</span><span class="w"> </span><span class="s1">'Arduino'</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="n">iex</span><span class="w"> </span><span class="s2">"mode </span><span class="si">$(</span><span class="bp">$_</span><span class="o">.</span><span class="nf">DeviceID</span><span class="si">)</span><span class="s2">: baud=9600 parity=n data=8 stop=1"</span><span class="p">;</span><span class="w"> </span><span class="nv">$port</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">new-Object</span><span class="w"> </span><span class="nx">System.IO.Ports.SerialPort</span><span class="w"> </span><span class="bp">$_</span><span class="o">.</span><span class="nf">DeviceID</span><span class="p">,</span><span class="nx">9600</span><span class="p">,</span><span class="nx">None</span><span class="p">,</span><span class="nx">8</span><span class="p">,</span><span class="nx">One</span><span class="p">;</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">Open</span><span class="p">();</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">WriteLine</span><span class="p">(</span><span class="bp">$_</span><span class="o">.</span><span class="nf">DeviceID</span><span class="p">);</span><span class="w"> </span><span class="nv">$port</span><span class="o">.</span><span class="nf">Close</span><span class="p">()}}</span><span class="w">
</span></code></pre></div></div>

<p>Therefore, the moment the AirCmd Tx transceiver is connected to the target device, it opens a PowerShell window, executes the port sensing commands, and stores the port information in memory. This stored information can later be used to redirect and write back the output of the PowerShell commands that were executed.</p>

<h3 id="powershell">PowerShell</h3>

<p>The PowerShell program executed by the AirCmd Tx transceiver stands as a highly versatile and powerful tool, serving a multitude of purposes across system administration, automation, and scripting tasks. An attacker may leverage PowerShell to bypass security measures like antivirus software due to its inherent trust within Windows environments. Its capability to execute code and scripts can be used to run malicious code or establish persistence on a system. Moreover, the extensive system access it offers can be exploited to extract sensitive data or manipulate system settings. The attacker might also harness PowerShell’s remote management features to control the compromised system or spread malware across a network. The combination of PowerShell’s versatility and power, along with its legitimate status, renders it a notable vector for attackers seeking to exploit vulnerabilities and exfiltrate data from the target system.</p>

<h2 id="demo">Demo</h2>

<div class="embed-youtube">
  <iframe src="https://www.youtube-nocookie.com/embed/lDhl9ivTlL0" title="Embedded YouTube video" loading="lazy" referrerpolicy="strict-origin-when-cross-origin" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen="">
  </iframe>
</div>

<p>The presented proof of concept has its limitations, reflecting certain imperfections inherent in the implementation phase. A more refined design could potentially address these limitations. One notable issue is AirCmd’s susceptibility to the constraints posed by LoRa technology, particularly in terms of data packet size and throughput, given its reliance on LoRa for data transfers. However, it’s crucial to remember the primary objective behind the creation of AirCmd: to demonstrate the feasibility of establishing networkless shells, and to shed light on how air-gapped systems might be compromised. Through this endeavor, I aimed to provide a tangible illustration of how these security challenges can manifest, while also offering a platform for exploring potential solutions or further investigations into overcoming the identified limitations.</p>

<p>The source code for this proof of concept is available <a href="https://github.com/Malware-Mystic/AirCmd/">here</a>.</p>]]></content><author><name>Nikhil Raghavendra</name></author><category term="Hardware &amp; Security" /><summary type="html"><![CDATA[A shell, in computing, refers to a user interface that provides access to an operating system’s services. It acts as an intermediary between the user and the operating system, allowing the user to execute commands, run scripts, and manage files and processes. Shells can be graphical (GUI) or command-line (CLI) based, with the latter being text-driven, where users input commands in text form and receive text responses. In a cybersecurity context, a shell refers to a command interface that attackers establish on a compromised system to remotely execute commands and control the system, often achieved through scripts or programs designed to exploit vulnerabilities in the target system. When discussing shells in the context of network security, two common types are mentioned: bind shells and reverse shells. Establishing a shell (both bind shell and reverse shell) necessitates a network connection between the attacker’s system and the target system. This network connection enables the transmission of commands and the retrieval of responses, making control and interaction with the remote system possible. However, in the case of air-gapped networks and isolated computing systems, establishing such a network connection is extremely challenging, almost reaching the realm of impossibility, due to the inherent design of these setups to remain disconnected from other networks as a security measure. The fundamental requirement of a network connection for establishing a shell presents a nearly insurmountable hurdle in the context of air-gapped networks and isolated computing systems. AirCmd AirCmd (pronounced “air command”) was conceived as a set of hardware and software toolkits that enable a penetration tester to set up a connectionless, networkless shell targeting an air-gapped network/isolated computing system. AirCmd allows a penetration tester to enjoy a bidirectional, shell-like command and control interface over LoRa. Given the long-range capabilities of LoRa, the penetration tester could be as far away as 5 km in urban areas, and up to 15 kilometers or more in rural areas (whithin line of sight) from the target. AirCmd doesn’t introduce any malware or malicious code into the target system, allowing it to evade standard anti-virus software and operate stealthily. AirCmd essentially operates as a “remote keyboard”, typing and executing commands on our behalf, albeit from a considerable distance. This proof of concept was conceived with several assumptions in mind: You have physical access to a computing device within the air-gapped network or isolated computing system. The computing device has a USB port that allows for the connection of peripherals via USB. The facility housing the air-gapped network or isolated computing system is not radio-hardened. PowerShell is available and usable on the target device. AirCmd is designed to allow you to input commands on your device seamlessly. Once a command is entered, the system promptly transmits it via LoRa to the transceiver hardware connected to the USB port of the target machine. The hardware then efficiently executes the transmitted command through PowerShell on the target machine. Following this execution, the output generated from the command is captured and swiftly transmitted back over LoRa to the transceiver hardware on your device. This mechanism empowers AirCmd to facilitate the remote execution of commands without the need for a network connection, thereby providing a secure, isolated layer of control over the targeted machine. Process flow of AirCmd Ax and Tx AirCmd consists of two segments: the attacker side, dubbed AirCmd Ax, and the target side, referred to as AirCmd Tx. The former houses two primary components: the AirCmd Shell and the AirCmd Ax transceiver. For this proof of concept, I utilized the LILYGO® LoRa32 V2.1_1.6. On the other hand, AirCmd Tx includes an AirCmd Tx transceiver, which integrates a Diymore BEETLE BadUSB Micro ATMEGA32U4-AU Development Expansion Module Board alongside an RFM95W-915S2 LoRa module. This segment also employs a PowerShell interface, which is instrumental for port sensing and command execution. AirCmd Shell The AirCmd Shell The AirCmd Shell is essentially a Python script that facilitates a connection for the attacker to the AirCmd Ax transceiver via USB serial. Upon issuing a command within the AirCmd Shell, the script encodes the command into UTF-8 format and transmits the data to the USB port to which the AirCmd Ax transceiver is connected. It then awaits a response from the AirCmd Ax transceiver, which is triggered upon receiving a transmission from the AirCmd Tx transceiver. AirCmd Ax Transceiver AirCmd Ax and Tx Transceiver Hardware The AirCmd Ax transceiver monitors the USB port it’s connected to for serial data. Upon receiving the data, it transmits it to the AirCmd Tx transceiver via LoRa. Following this, it awaits a response from the AirCmd Tx transceiver. Once a response is received, it relays this data back to the AirCmd Shell via the serial connection. AirCmd Tx Transceiver The AirCmd Tx transceiver is the most crucial component of AirCmd. It acts as a gateway for receiving the payload, injecting it into the target via PowerShell, and retrieving the data from the interface before transmitting it to the AirCmd Ax transceiver. Injecting and executing a PowerShell payload on the target device is straightforward and can be easily accomplished using Arduino’s Keyboard library. However, our objective extends beyond mere execution; we aim to execute the command and retrieve its output. For instance, if we wish to view the contents of a file, merely executing the command wouldn’t suffice. Although the command will display the file’s contents, there isn’t a direct method to view these contents via the AirCmd Shell. Therefore, we employ a series of PowerShell commands to execute the original command, pipe its output, consolidate the multiple lines of output into a single line separated by a designated separator, and then transmit this data to the serial port to which our AirCmd Tx transceiver is connected. $lines = [CMD] | % { $_.ToString() }; $joinedLines = $lines -join '?'; $port = new-Object System.IO.Ports.SerialPort [PORT], 9600, None, 8, One; $port.Open(); $port.WriteLine($joinedLines); $port.Close(); The command [CMD] is issued within the AirCmd Shell and is set to execute, with its output being directed to the serial [PORT] to which the AirCmd Tx transceiver is connected. However, a question arises: how does the AirCmd Tx transceiver discern the specific [PORT] to which it’s connected? To address this, we employ a series of PowerShell commands designed for port sensing. Get-WmiObject -Class Win32_SerialPort | % {if ($_.Description -match 'Arduino') {iex "mode $($_.DeviceID): baud=9600 parity=n data=8 stop=1"; $port = new-Object System.IO.Ports.SerialPort $_.DeviceID,9600,None,8,One; $port.Open(); $port.WriteLine($_.DeviceID); $port.Close()}} Therefore, the moment the AirCmd Tx transceiver is connected to the target device, it opens a PowerShell window, executes the port sensing commands, and stores the port information in memory. This stored information can later be used to redirect and write back the output of the PowerShell commands that were executed. PowerShell The PowerShell program executed by the AirCmd Tx transceiver stands as a highly versatile and powerful tool, serving a multitude of purposes across system administration, automation, and scripting tasks. An attacker may leverage PowerShell to bypass security measures like antivirus software due to its inherent trust within Windows environments. Its capability to execute code and scripts can be used to run malicious code or establish persistence on a system. Moreover, the extensive system access it offers can be exploited to extract sensitive data or manipulate system settings. The attacker might also harness PowerShell’s remote management features to control the compromised system or spread malware across a network. The combination of PowerShell’s versatility and power, along with its legitimate status, renders it a notable vector for attackers seeking to exploit vulnerabilities and exfiltrate data from the target system. Demo The presented proof of concept has its limitations, reflecting certain imperfections inherent in the implementation phase. A more refined design could potentially address these limitations. One notable issue is AirCmd’s susceptibility to the constraints posed by LoRa technology, particularly in terms of data packet size and throughput, given its reliance on LoRa for data transfers. However, it’s crucial to remember the primary objective behind the creation of AirCmd: to demonstrate the feasibility of establishing networkless shells, and to shed light on how air-gapped systems might be compromised. Through this endeavor, I aimed to provide a tangible illustration of how these security challenges can manifest, while also offering a platform for exploring potential solutions or further investigations into overcoming the identified limitations. The source code for this proof of concept is available here.]]></summary></entry><entry><title type="html">Hacking a Live Streaming App</title><link href="https://nikhilr.io/posts/Hacking-a-Live-Streaming-App/" rel="alternate" type="text/html" title="Hacking a Live Streaming App" /><published>2023-07-28T00:00:00+08:00</published><updated>2023-07-28T00:00:00+08:00</updated><id>https://nikhilr.io/posts/Hacking-a-Live-Streaming-App</id><content type="html" xml:base="https://nikhilr.io/posts/Hacking-a-Live-Streaming-App/"><![CDATA[<blockquote>
  <p>This blog post, originally written in 2019, was migrated from my old blog site.</p>
</blockquote>

<p>The ICC Cricket World Cup concluded a few weeks ago, and it was watched by approximately 2.6 billion people worldwide. This staggering viewership makes it the most-watched cricket competition as of 2019. The advent of on-demand mobile media applications such as Netflix and YouTube is increasingly attracting people like me away from traditional television sets with cable or fiber connections, shifting the dependence onto native or web apps for media content.</p>

<p>Cricket is typically a “seasonal” sport. High-level cricket is almost always played outdoors, on uncovered pitches, and play is halted by rain. The seasons in each country are arranged to coincide with the driest months of the year. Factors such as hours of daylight and temperature also influence the scheduling in some countries. For example, in England, the winter days are too short and often too dark, and the temperature too low for cricket to be a viable option. Thus, cricket is typically played outdoors, but in the UK, the sport is played indoors once the season concludes. When referring to cricket seasons, the convention is to use a single year for a northern hemisphere summer season and a dashed pair of years to indicate a southern hemisphere summer. In tropical regions, cricket can be played throughout the year. In the United Kingdom, the cricket season begins in mid-April and ends in September, whereas in Australia, the season kicks off in October and concludes in February or March.</p>

<p>Most people only follow a handful of cricket teams. Consequently, having a cable or fiber TV connection solely for a sports channel might not be cost-effective, particularly for those wishing to maximize their TV subscription by watching a seasonal sport like cricket. As a result, a significant number of cricket fans have migrated to live streaming platforms on mobile or web apps, most of which are either free, ad-supported, or require a small subscription fee.</p>

<p>After exploring the Google Play Store, I found a reputable Android application with positive reviews to live stream the ICC Cricket World Cup tournament. I downloaded the app onto my Samsung tablet and began watching the tournament. However, watching a game that lasts for hours on a tablet can be strenuous on both my hands and eyes. Therefore, I decided to watch the match live on TV, and here’s how I achieved it.</p>

<h2 id="setting-up">Setting Up</h2>

<p>I utilized an online APK downloader to retrieve the app’s APK file from the Google Play Store. I am not going to recommend any specific online APK downloader because that boils down to personal preference, and it’s fairly easy to find reputable ones online. After downloading the APK file, I launched Android Studio, opened the AVD manager by navigating to <code class="language-plaintext highlighter-rouge">Tools &gt; AVD Manager</code>, and started a Pixel 2 XL (Android 7.1.1) AVD in the emulator. I installed the app on the AVD by dragging and dropping the APK file onto the emulator screen.</p>

<p><img src="/assets/img/hack1.png" alt="img" />
<em>The Pixel 2 XL Android Virtual Device (AVD) in the emulator</em></p>

<p>I launched the app by locating and clicking on it in the AVD’s app menu. As is common with ad-monetized apps, I was immediately bombarded with advertisements upon launching the app; the app’s media player screen was also cluttered with banner ads. I needed to determine the URL from which the media player was streaming the live broadcast. I closed my browser, deactivated my VPN, and launched Wireshark to monitor some packets.</p>

<h2 id="packet-analysis">Packet Analysis</h2>

<p>Wireshark is a network analysis tool that captures packets in real time and presents them in a format that is easy to understand. You can download Wireshark for your operating system from the <a href="https://www.wireshark.org/#download">official Wireshark website</a>. I connect to the Internet using my <code class="language-plaintext highlighter-rouge">WiFi: en0</code> interface, which is also the interface the emulator uses to stream the live broadcast.</p>

<p><img src="/assets/img/hack2.png" alt="img" />
<em>I aim to capture packets from the WiFi: en0 interface on my Mac</em></p>

<p>To start capturing packets from the <code class="language-plaintext highlighter-rouge">WiFi: en0</code> interface, I clicked the blue-colored shark fin located in the upper left-hand corner of the menu (prior to capturing packets, I ensured that my browser was closed, my VPN was turned off, and any running application was quit). I captured packets from the <code class="language-plaintext highlighter-rouge">WiFi: en0</code> interface for about 30 seconds, allowing Wireshark to analyze the packets being sent and received by the interface more accurately.</p>

<p><img src="/assets/img/hack3.png" alt="img" />
<em>The app performs two different GET requests that is of interest to us</em></p>

<p>The app appears to perform a GET request to retrieve two distinct files with different extensions: <code class="language-plaintext highlighter-rouge">.ts</code> and <code class="language-plaintext highlighter-rouge">.m3u8</code>. The MPEG transport stream (with file extensions such as <code class="language-plaintext highlighter-rouge">.ts</code>, <code class="language-plaintext highlighter-rouge">.tsv</code>, <code class="language-plaintext highlighter-rouge">.tsa</code>) is a standard digital container format used for transmitting and storing audio, video, and Program and System Information Protocol (PSIP) data. It is employed in broadcasting systems such as DVB, ATSC, and IPTV. The app I’m using leverages Internet Protocol Television (IPTV) to transmit television content over Internet Protocol (IP) networks. This differs from traditional terrestrial, satellite, and cable television formats. Unlike downloaded media, IPTV allows for continuous streaming of source media. As a result, a client media player can start playing the content (such as a TV channel) almost instantly. This phenomenon is known as streaming media.</p>

<p>M3U (MP3 URL or Moving Picture Experts Group Audio Layer 3 Uniform Resource Locator in full; file extensions: <code class="language-plaintext highlighter-rouge">.m3u</code>, <code class="language-plaintext highlighter-rouge">.m3u8</code>) is a computer file format for a multimedia playlist. A common use of the M3U file format is creating a single-entry playlist file that points to a stream on the Internet. The created file provides easy access to that stream and is often utilized in downloads from a website, in emailing, and in listening to Internet radio.</p>

<p>Although it was initially designed for audio files, like MP3, it’s commonly used to point media players to audio and video sources, including online ones. Fraunhofer originally developed M3U for use with their Winplay3 software, but numerous media player and software application developers quickly adopted the standard.</p>

<p>M3U8 is the Unicode version of M3U, which utilizes UTF-8-encoded characters. M3U8 files are the foundation for the HTTP Live Streaming (HLS) format initially developed by Apple to stream video and radio to iOS devices. It is now a popular format for Dynamic Adaptive Streaming over HTTP (DASH) in general.</p>

<p>The app retrieves a <code class="language-plaintext highlighter-rouge">.m3u8</code> file, the Unicode version of M3U, commonly used to direct media players to audio and video sources, including online sources, to live stream the broadcast. It then retrieves a <code class="language-plaintext highlighter-rouge">.ts</code> file, a standard digital container format that encapsulates packetized elementary streams, equipped with error correction and synchronization pattern features that maintain transmission integrity when the communication channel for the stream is compromised.</p>

<p>We need to discover the URL from which the <code class="language-plaintext highlighter-rouge">.m3u8</code> file is retrieved to access the <code class="language-plaintext highlighter-rouge">.m3u8</code> file via the browser. Luckily, Wireshark has captured and analyzed the packet for us, so we can easily retrieve the URL.</p>

<p><img src="/assets/img/hack4.png" alt="img" />
<em>The full request URL from which we can stream the live broadcast</em></p>

<p>The complete request URL from which the <code class="language-plaintext highlighter-rouge">.m3u8</code> file is retrieved can be found under the Hypertext Transfer Protocol section. I copied and pasted the URL into Safari’s search box, and I was able to view the live broadcast instantly. To AirPlay the live stream to my TV, I clicked on the AirPlay button located in Safari’s default video player. Now, I was able to watch the match live on my TV!</p>

<p><img src="/assets/img/hack5.png" alt="img" /></p>

<h2 id="osint">OSINT</h2>

<p>Wanting to understand more about the app’s back-end provider, I recognized that it’s nearly impossible for an independent developer to construct and manage his own server infrastructure capable of supporting hundreds of thousands of simultaneous live stream views. Therefore, I executed a <code class="language-plaintext highlighter-rouge">whois</code> lookup on the server’s IP address, and here’s what I discovered.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[TRUNCATED]
inetnum:        109.169.0.0 - 109.169.95.255
netname:        UK-RAPIDSWITCH-20091102
country:        GB
org:            ORG-RL20-RIPE
admin-c:        AR6363-RIPE
tech-c:         AR6363-RIPE
status:         ALLOCATED PA
mnt-by:         RIPE-NCC-HM-MNT
mnt-by:         RAPIDSWITCH-MNT
mnt-routes:     RAPIDSWITCH-MNT
created:        2010-02-11T09:11:40Z
last-modified:  2017-03-24T16:04:24Z
source:         RIPE # Filtered

organisation:   ORG-RL20-RIPE
org-name:       IOMART HOSTING LIMITED
org-type:       LIR
address:        Spectrum House, Clivemont Road
address:        SL6 7FW
address:        Maidenhead
address:        UNITED KINGDOM
phone:          +441753471040
fax-no:         +441753471049
admin-c:        RC6613-RIPE
admin-c:        DB16530-RIPE
admin-c:        RM1358-RIPE
admin-c:        SMC74-RIPE
admin-c:        AR6363-RIPE
mnt-ref:        RAPIDSWITCH-MNT
mnt-ref:        RIPE-NCC-HM-MNT
mnt-by:         RIPE-NCC-HM-MNT
mnt-by:         RAPIDSWITCH-MNT
abuse-c:        AR12896-RIPE
created:        2005-09-26T12:37:33Z
last-modified:  2018-12-11T14:48:04Z
source:         RIPE # Filtered

person:         Abuse Robot
address:        iomart Hosting Ltd t/a RapidSwitch
address:        Spectrum House
address:        Clivemont Road
address:        Maidenhead
address:        SL6 7FW
phone:          +44 (0)1753 471 040
remarks:        ******************************************************
remarks:        * ABUSE REPORTS                                      *
remarks:        * https://myservers.rapidswitch.com/reportabuse.aspx *
remarks:        ******************************************************
nic-hdl:        AR6363-RIPE
mnt-by:         RAPIDSWITCH-MNT
created:        2007-02-11T09:38:19Z
last-modified:  2017-10-30T21:53:52Z
source:         RIPE # Filtered
[TRUNCATED]
</code></pre></div></div>

<p>The server broadcasting the live stream is owned by iomart Hosting Limited, which trades as RapidSwitch. Their data center is situated in Spectrum House, Maidenhead SL6 7FW, UK, as confirmed by Google Maps.</p>

<p><img src="/assets/img/hack6.png" alt="img" />
<em>RapidSwitch’s data center is located in Spectrum House, United Kingdom</em></p>

<p>We can also learn more about the frameworks/technologies that the app uses to stream the live broadcast by disassembling the app’s APK file. I’ll conclude here for now because I need to watch the rest of the match. Goodbye!</p>]]></content><author><name>Nikhil Raghavendra</name></author><category term="Security" /><summary type="html"><![CDATA[This blog post, originally written in 2019, was migrated from my old blog site. The ICC Cricket World Cup concluded a few weeks ago, and it was watched by approximately 2.6 billion people worldwide. This staggering viewership makes it the most-watched cricket competition as of 2019. The advent of on-demand mobile media applications such as Netflix and YouTube is increasingly attracting people like me away from traditional television sets with cable or fiber connections, shifting the dependence onto native or web apps for media content. Cricket is typically a “seasonal” sport. High-level cricket is almost always played outdoors, on uncovered pitches, and play is halted by rain. The seasons in each country are arranged to coincide with the driest months of the year. Factors such as hours of daylight and temperature also influence the scheduling in some countries. For example, in England, the winter days are too short and often too dark, and the temperature too low for cricket to be a viable option. Thus, cricket is typically played outdoors, but in the UK, the sport is played indoors once the season concludes. When referring to cricket seasons, the convention is to use a single year for a northern hemisphere summer season and a dashed pair of years to indicate a southern hemisphere summer. In tropical regions, cricket can be played throughout the year. In the United Kingdom, the cricket season begins in mid-April and ends in September, whereas in Australia, the season kicks off in October and concludes in February or March. Most people only follow a handful of cricket teams. Consequently, having a cable or fiber TV connection solely for a sports channel might not be cost-effective, particularly for those wishing to maximize their TV subscription by watching a seasonal sport like cricket. As a result, a significant number of cricket fans have migrated to live streaming platforms on mobile or web apps, most of which are either free, ad-supported, or require a small subscription fee. After exploring the Google Play Store, I found a reputable Android application with positive reviews to live stream the ICC Cricket World Cup tournament. I downloaded the app onto my Samsung tablet and began watching the tournament. However, watching a game that lasts for hours on a tablet can be strenuous on both my hands and eyes. Therefore, I decided to watch the match live on TV, and here’s how I achieved it. Setting Up I utilized an online APK downloader to retrieve the app’s APK file from the Google Play Store. I am not going to recommend any specific online APK downloader because that boils down to personal preference, and it’s fairly easy to find reputable ones online. After downloading the APK file, I launched Android Studio, opened the AVD manager by navigating to Tools &gt; AVD Manager, and started a Pixel 2 XL (Android 7.1.1) AVD in the emulator. I installed the app on the AVD by dragging and dropping the APK file onto the emulator screen. The Pixel 2 XL Android Virtual Device (AVD) in the emulator I launched the app by locating and clicking on it in the AVD’s app menu. As is common with ad-monetized apps, I was immediately bombarded with advertisements upon launching the app; the app’s media player screen was also cluttered with banner ads. I needed to determine the URL from which the media player was streaming the live broadcast. I closed my browser, deactivated my VPN, and launched Wireshark to monitor some packets. Packet Analysis Wireshark is a network analysis tool that captures packets in real time and presents them in a format that is easy to understand. You can download Wireshark for your operating system from the official Wireshark website. I connect to the Internet using my WiFi: en0 interface, which is also the interface the emulator uses to stream the live broadcast. I aim to capture packets from the WiFi: en0 interface on my Mac To start capturing packets from the WiFi: en0 interface, I clicked the blue-colored shark fin located in the upper left-hand corner of the menu (prior to capturing packets, I ensured that my browser was closed, my VPN was turned off, and any running application was quit). I captured packets from the WiFi: en0 interface for about 30 seconds, allowing Wireshark to analyze the packets being sent and received by the interface more accurately. The app performs two different GET requests that is of interest to us The app appears to perform a GET request to retrieve two distinct files with different extensions: .ts and .m3u8. The MPEG transport stream (with file extensions such as .ts, .tsv, .tsa) is a standard digital container format used for transmitting and storing audio, video, and Program and System Information Protocol (PSIP) data. It is employed in broadcasting systems such as DVB, ATSC, and IPTV. The app I’m using leverages Internet Protocol Television (IPTV) to transmit television content over Internet Protocol (IP) networks. This differs from traditional terrestrial, satellite, and cable television formats. Unlike downloaded media, IPTV allows for continuous streaming of source media. As a result, a client media player can start playing the content (such as a TV channel) almost instantly. This phenomenon is known as streaming media. M3U (MP3 URL or Moving Picture Experts Group Audio Layer 3 Uniform Resource Locator in full; file extensions: .m3u, .m3u8) is a computer file format for a multimedia playlist. A common use of the M3U file format is creating a single-entry playlist file that points to a stream on the Internet. The created file provides easy access to that stream and is often utilized in downloads from a website, in emailing, and in listening to Internet radio. Although it was initially designed for audio files, like MP3, it’s commonly used to point media players to audio and video sources, including online ones. Fraunhofer originally developed M3U for use with their Winplay3 software, but numerous media player and software application developers quickly adopted the standard. M3U8 is the Unicode version of M3U, which utilizes UTF-8-encoded characters. M3U8 files are the foundation for the HTTP Live Streaming (HLS) format initially developed by Apple to stream video and radio to iOS devices. It is now a popular format for Dynamic Adaptive Streaming over HTTP (DASH) in general. The app retrieves a .m3u8 file, the Unicode version of M3U, commonly used to direct media players to audio and video sources, including online sources, to live stream the broadcast. It then retrieves a .ts file, a standard digital container format that encapsulates packetized elementary streams, equipped with error correction and synchronization pattern features that maintain transmission integrity when the communication channel for the stream is compromised. We need to discover the URL from which the .m3u8 file is retrieved to access the .m3u8 file via the browser. Luckily, Wireshark has captured and analyzed the packet for us, so we can easily retrieve the URL. The full request URL from which we can stream the live broadcast The complete request URL from which the .m3u8 file is retrieved can be found under the Hypertext Transfer Protocol section. I copied and pasted the URL into Safari’s search box, and I was able to view the live broadcast instantly. To AirPlay the live stream to my TV, I clicked on the AirPlay button located in Safari’s default video player. Now, I was able to watch the match live on my TV! OSINT Wanting to understand more about the app’s back-end provider, I recognized that it’s nearly impossible for an independent developer to construct and manage his own server infrastructure capable of supporting hundreds of thousands of simultaneous live stream views. Therefore, I executed a whois lookup on the server’s IP address, and here’s what I discovered. [TRUNCATED] inetnum: 109.169.0.0 - 109.169.95.255 netname: UK-RAPIDSWITCH-20091102 country: GB org: ORG-RL20-RIPE admin-c: AR6363-RIPE tech-c: AR6363-RIPE status: ALLOCATED PA mnt-by: RIPE-NCC-HM-MNT mnt-by: RAPIDSWITCH-MNT mnt-routes: RAPIDSWITCH-MNT created: 2010-02-11T09:11:40Z last-modified: 2017-03-24T16:04:24Z source: RIPE # Filtered organisation: ORG-RL20-RIPE org-name: IOMART HOSTING LIMITED org-type: LIR address: Spectrum House, Clivemont Road address: SL6 7FW address: Maidenhead address: UNITED KINGDOM phone: +441753471040 fax-no: +441753471049 admin-c: RC6613-RIPE admin-c: DB16530-RIPE admin-c: RM1358-RIPE admin-c: SMC74-RIPE admin-c: AR6363-RIPE mnt-ref: RAPIDSWITCH-MNT mnt-ref: RIPE-NCC-HM-MNT mnt-by: RIPE-NCC-HM-MNT mnt-by: RAPIDSWITCH-MNT abuse-c: AR12896-RIPE created: 2005-09-26T12:37:33Z last-modified: 2018-12-11T14:48:04Z source: RIPE # Filtered person: Abuse Robot address: iomart Hosting Ltd t/a RapidSwitch address: Spectrum House address: Clivemont Road address: Maidenhead address: SL6 7FW phone: +44 (0)1753 471 040 remarks: ****************************************************** remarks: * ABUSE REPORTS * remarks: * https://myservers.rapidswitch.com/reportabuse.aspx * remarks: ****************************************************** nic-hdl: AR6363-RIPE mnt-by: RAPIDSWITCH-MNT created: 2007-02-11T09:38:19Z last-modified: 2017-10-30T21:53:52Z source: RIPE # Filtered [TRUNCATED] The server broadcasting the live stream is owned by iomart Hosting Limited, which trades as RapidSwitch. Their data center is situated in Spectrum House, Maidenhead SL6 7FW, UK, as confirmed by Google Maps. RapidSwitch’s data center is located in Spectrum House, United Kingdom We can also learn more about the frameworks/technologies that the app uses to stream the live broadcast by disassembling the app’s APK file. I’ll conclude here for now because I need to watch the rest of the match. Goodbye!]]></summary></entry></feed>