Rehumanize: Ambient AI Content Detection as Infrastructure for Synthetic Media Transparency
Rehumanize
Ambient AI Content Detection as Infrastructure for Synthetic Media Transparency
Problem Statement
Every day, billions of people turn to social media not just for entertainment, but to understand the world - to gauge what others think, to form opinions, and to consume news (DataReportal, 2025). This act of social calibration has always depended on an implicit assumption: that the voices in the feed are human. In the age of AI, that assumption is no longer safe. When an information environment is synthetically populated, misinformation can take on a new form: the apparent social evidence around a claim may itself be manufactured, even when no individual sentence is factually false.
This is a new dimension of misinformation in the age of AI. The problem is no longer limited to the traditional binary of true versus false information. It has acquired added dimensions: authenticity of origin, manufactured consensus, and the synthetic population of information environments. A claim can be true or false, but the apparent human voice, experience, or consensus surrounding it can also be authentic or synthetic. Rehumanize addresses that added layer because misinformation can now be amplified not only through false claims, but through false signals of human agreement.
And yet - crucially, synthetic does not automatically mean false. A synthetic but accurate summary of verified election results from a news outlet is genuinely useful. A synthetic weather forecast harms no one. But a synthetic opinion designed to look like grassroots testimony can manufacture the appearance of public support. A synthetic product review can create a false impression of genuine experience. The harm is not that AI-generated content exists; it is that people scrolling their feeds cannot currently distinguish its synthetic origin from human expression. In the misinformation ecosystem, that missing context matters because users may give synthetic voices the same social weight they would give genuine ones.
AI-generated text is arriving on social media at a scale and pace that outstrips existing awareness. Research analysing around 2.4 million posts across Medium, Quora, and Reddit has documented the growing presence of AI-generated text on social platforms (Sun et al., 2025). These are not fringe environments. They are places where ordinary people seek information, advice, and social cues. As synthetic content becomes harder to distinguish from human expression, it creates a new misinformation risk: users can be influenced by the perceived presence of human voices and consensus even when the underlying content is not false in a conventional fact-checking sense.
The conventional response to misinformation focuses on false facts: claims that can be checked, corrected, and labelled. That remains necessary, but it is no longer sufficient. In the age of AI, the misinformation problem has acquired another dimension: users also need to know whether the apparent human source of a message is actually human. Rehumanize therefore asks a complementary question: "is this human or synthetic?" These are not the same question. A fact-checker can verify whether a claim is accurate, but a thousand AI-generated comments expressing the same political opinion can manufacture the appearance of human consensus without presenting one discrete factual claim. That synthetic social signal can become part of the misinformation mechanism itself.
As Woolley (2023) documented in Manufacturing Consensus, automated systems can create systematic illusions of agreement that shape public discourse. In the AI era, that mechanism becomes a direct extension of misinformation risk: when AI systems generate posts, comments, and reactions at scale, they can manufacture the appearance of human consensus where none exists. The consensus is not organic - it is engineered.
This mechanism has already been documented in practice. In late 2023, BBC and Clemson University investigators revealed DCWeekly.org to be a Russian-linked influence operation that adopted GPT-3 in September 2023 to rewrite and publish pro-Kremlin content at scale. A peer-reviewed analysis of 22,889 articles confirmed that AI adoption measurably increased the operation's output volume and topical breadth, without any reduction in persuasiveness (Wack et al., 2025). Separately, a German Federal Foreign Office investigation identified the Doppelgänger operation: a coordinated pro-Russian campaign between December 2023 and January 2024 that generated over one million German-language posts via more than 50,000 fake accounts on X, designed to undermine Germany's support for Ukraine by spreading narratives that the government prioritised foreign aid over domestic needs (Saab, 2024). These are not isolated incidents - they represent a documented, cross-jurisdictional pattern of AI-enabled information manipulation operating across different formats (articles and social media posts), different countries, and different platforms.
Three Dimensions of Harm
Trust calibration. Research confirms that repeated exposure to content - including false content - increases belief through the illusory truth effect, and that this effect grows with additional exposures (Udry & Barber, 2024). Repetition builds belief. In misinformation campaigns, synthetic repetition adds another layer: the repeated voices can themselves be manufactured. Synthetic content is artificially cheap to produce at scale, allowing one actor to create the appearance of many independent people reinforcing the same narrative. Rehumanize therefore treats synthetic origin as information relevant to how much social weight a user should give repeated claims and opinions.
Persuasion at scale. Large language models can generate politically persuasive content quickly and at scale, with research showing that LLM-generated messages can measurably change attitudes across policy issues (Hackenburg & Margetts, 2024; Bai et al., 2025). This matters for misinformation because a synthetic voice can now deliver persuasive narratives without the cost of producing equivalent volumes of human-authored messaging. The risk is not that every persuasive AI message is misinformation; it is that the same capability can be used to amplify misleading narratives, simulate grassroots sentiment, or overwhelm genuine voices.
Epistemic autonomy. When the social environment people use to form views is substantially synthetic, the process of forming genuine independent judgements can be compromised at its root (Jakesch, Hancock & Naaman, 2023). This is directly aligned with Media and Information Literacy: critical evaluation requires knowing not only whether information is accurate, but what kind of information environment produced it. If users cannot see that the voices surrounding a claim are synthetic, misinformation can operate through deception about who is speaking and how much apparent agreement actually exists.
Why Existing Solutions Fail
Platform-level disclosure frameworks exist but are structurally insufficient for the misinformation problem created by synthetic social signals.
The EU AI Act's Article 50 transparency obligations apply from 2 August 2026 and include machine-readable marking of AI-generated or manipulated content, alongside clear labelling requirements for certain AI-generated text and deepfakes. This is an important regulatory layer because it recognises that synthetic origin can create risks of deception and manipulation. Technical watermarking is also advancing: Google DeepMind's SynthID-Text is a production-ready watermarking approach that embeds a statistical signal during text generation (Dathathri et al., 2024). These developments strengthen transparency, but they do not remove the misinformation problem entirely. Watermarking depends on participating generation systems, while bad-faith actors can use other models or alter content.
Fact-checking tools address individual claims after publication, but synthetic misinformation can also operate through the apparent identity and volume of the voices carrying those claims. Media literacy education requires deliberate critical engagement - curriculum, workshops, institutional access, and personal motivation - while social media is often consumed passively and at speed. Expecting a user scrolling Threads at midnight to stop and determine both whether a claim is true and whether the surrounding voices are synthetic is a fundamental mismatch between the mode of consumption and the mode of intervention. Third-party ambient detection therefore adds a transparency layer to the misinformation response: it surfaces synthetic origin while the user is consuming the information environment.
This is the gap Rehumanize addresses: not replacing fact-checking, but adding visibility into the synthetic voices and consensus signals through which misinformation can now spread.
Proposed Solution Overview
Rehumanize is an ambient synthetic content transparency tool for social media. It continuously scans user feeds and quietly surfaces a visual indicator - an amber border - when content is likely AI-generated, requiring no deliberate action from the user. The purpose is not to declare a post misinformation. It is to expose a new piece of context that matters when misinformation is created, amplified, or made to appear socially supported.
The technical flow is straightforward: as the user scrolls, the extension intercepts post text in real time, sends it to a transformer-based AI detection model, receives a probability score, and - if the score exceeds a calibrated threshold - renders the amber indicator on the post. The entire process happens without interrupting the scroll experience. The user does not need to suspect anything, copy anything, navigate anywhere, or wait for anything. The detection is ambient: it makes synthetic origin visible at the point where users are forming impressions, evaluating claims, and calibrating trust.
This complements rather than replaces existing detection and fact-checking tools. Paste-and-check services require a user to already suspect that a post is AI-generated, then copy the text, navigate to a separate website, paste it, and wait for results. That workflow is poorly matched to the way misinformation and synthetic social signals are encountered in fast-moving feeds. Rehumanize moves the transparency signal into the feed itself, where users can see it before the synthetic voice contributes to their judgement.
A working prototype is already deployed and live on the Chrome Web Store, operating across Threads, Reddit, and X.
A critical design principle: Rehumanize does not flag content as wrong. It flags content as likely not human-authored. The amber indicator is not a red flag saying "this is bad" - it is a transparency layer saying "this is likely synthetic; now you decide." A synthetic summary of verified news may be useful. A synthetic "personal story" designed to look like grassroots testimony can be part of a misinformation campaign because it manufactures the appearance of lived experience. A synthetic product review can create a deceptive endorsement. The tool does not make these distinctions for the user - it gives the user the information needed to make them. That is Media and Information Literacy applied to a new dimension of misinformation: not telling people what to think, but making the information environment more legible.
Rehumanize sits at the intersection of AI-driven detection and media literacy research. The detection layer uses transformer-based NLP models to identify synthetic content in real time. The research layer uses peer-reviewed tools from Cardiff NLP (EMNLP 2022; COLING 2022) to characterise flagged content by sentiment and topic, generating aggregate evidence about how synthetic material appears in real-world social feeds. This matters to misinformation research because it can help measure not only false claims, but the prevalence and characteristics of the synthetic voices that surround, repeat, and amplify them.
Implementation and Proof of Concept
The proof-of-concept demonstration shows the interface concept. Posts identified as likely AI-generated display the amber border indicator. Posts that are too short to analyse reliably display a separate badge. The interface requires no interaction - it operates as the user scrolls, adding a transparency signal directly to the environment in which misinformation and social consensus are encountered.
Research Pipeline
Beyond user-facing transparency, Rehumanize incorporates a research data pipeline designed to generate aggregate insights on synthetic content prevalence without storing any personal data.
When a post is flagged as likely AI-generated, an asynchronous secondary process runs: the post text is sent through Cardiff NLP's TweetNLP models - specifically cardiffnlp/twitter-roberta-base-sentiment-latest (RoBERTa-base, trained on approximately 124 million tweets, fine-tuned on the TweetEval benchmark, CC-BY-4.0 licensed, peer-reviewed at EMNLP 2022) for sentiment classification (positive, negative, neutral), and cardiffnlp/tweet-topic-21-multi (same underlying TimeLMs language model, fine-tuned on 11,267 annotated tweets for multi-label topic classification across 19 categories including politics, health, sports, and more, peer-reviewed at COLING 2022). The text is then immediately discarded. This process does not affect or delay the real-time detection response shown to the user.
What is logged per flagged post: platform, timestamp (day/hour bucketed - never precise), AI detection score, sentiment label with confidence score, and topic label with confidence score. No text is stored. No account identifiers. No user identifiers. Data collection activates only with explicit user consent via an informed opt-in during onboarding, and consent is revocable at any time.
The resulting research layer is designed to support aggregate analysis of synthetic content prevalence, topic distribution, and sentiment patterns without storing the underlying post text or user identifiers. This is valuable for misinformation research because it can reveal where synthetic voices are concentrated and how they interact with the information environments in which people form beliefs.
Expected Impact
Epistemic Impact: The Right to Know
The problem Rehumanize addresses is not that AI-generated content exists. It is that people navigating social media feeds have no reliable way of knowing which apparent voices are synthetic - and that invisibility can become part of the misinformation mechanism.
The immediate impact is epistemic rather than behavioural. Rehumanize provides information, not judgement. The amber indicator does not say "this is misinformation." It says "this was likely not written by a human."
A synthetic summary of election results from a news outlet? Probably useful - the user sees the indicator, recognises the AI assistance, and evaluates the content on its merits. A synthetic "personal story" in a political discussion, designed to look like grassroots testimony? That is manufactured consensus, and knowing it is synthetic fundamentally changes how it should be weighted. A synthetic product review? Deceptive - the user benefits from knowing the endorsement is not genuine. A synthetic weather update? Entirely harmless - no one is hurt by an AI writing a forecast.
The tool does not make these distinctions for the user - it enables the user to make them. This is the difference between a paternalistic intervention (telling people what to believe) and a Media and Information Literacy intervention (giving people information needed to evaluate misinformation and synthetic social signals for themselves). The goal is not lower trust - it is calibrated trust.
This aligns with the EU AI Act's Article 50 transparency framework, which applies from 2 August 2026 and requires clear labelling of certain AI-generated text and machine-readable marking of AI-generated or manipulated content. The principle is directly relevant to misinformation: users need visibility into the composition and provenance of the information environment before they can calibrate how much weight to give what they encounter. Rehumanize operationalises that transparency principle at the point of consumption.
UNESCO's Media and Information Literacy framework centres on citizens' ability to access, evaluate, and create information critically. Rehumanize directly enables the evaluate step by making synthetic origin visible - adding a capability that matters because misinformation can now involve both the truth of a claim and the authenticity of the voices presenting it.
Traditional MIL assumes a pipeline: teach people critical thinking skills, then expect them to apply those skills when consuming media. This pipeline has a structural bottleneck. It requires education programmes, curricula, institutional access, and personal motivation. It reaches the already-engaged and misses the most vulnerable. Ambient detection adds a complementary layer: instead of expecting users to recognise synthetic voices unaided, it makes likely synthetic origin visible by default. The literacy is embedded in the infrastructure itself.
Every scroll becomes a moment of awareness - not because the user is studying, but because the information environment has been made more legible. This is MIL that operates in the passive consumption mode where misinformation is encountered. It does not replace traditional media literacy education - it provides an ambient foundation on which that education can build.
Furthermore, the research data generated by Rehumanize - aggregate patterns of synthetic content prevalence, topic distribution, and sentiment across platforms - can contribute to academic MIL research on misinformation. It can help answer questions such as: how much of a typical feed is synthetic, how does that vary by platform and topic, and where do synthetic voices cluster around public-interest events? Rehumanize is simultaneously a user-facing MIL intervention and a research instrument for understanding a new dimension of the information environment.
Watermarking as a Complementary Layer
The landscape of synthetic content transparency is evolving rapidly. Google DeepMind's SynthID-Text demonstrates that statistical watermarking can provide a machine-detectable signal for AI-generated text while preserving output quality (Dathathri et al., 2024). The EU AI Act's transparency framework now reinforces this direction by requiring machine-readable marking of certain AI-generated or manipulated content. These measures are important, but they remain provider-side signals and therefore do not eliminate the need for user-side detection when content comes from non-participating or non-compliant systems.
This is a significant development. However, provider-side watermarking and user-side ambient detection serve complementary functions. Watermarking can provide strong provenance signals when generation systems participate; ambient detection can surface likely synthetic content in the feed even when no reliable watermark is available. Together they form complementary layers of synthetic-content transparency within the broader misinformation response.
Research Infrastructure
The long-term vision for Rehumanize extends beyond individual users. The research pipeline - using peer-reviewed Cardiff NLP models to classify flagged content by sentiment and topic, logging only aggregate metadata with no text or identifiers stored - is designed to generate naturalistic aggregate evidence about synthetic content across social platforms. That evidence can help misinformation researchers study not only false claims, but the synthetic voices and consensus signals that may surround them.
This data enables research questions about the misinformation implications of synthetic social environments:
Election integrity monitoring. How does synthetic content cluster around political events? Are there detectable spikes in AI-generated political content around election periods? Which political topics attract the most synthetic content? Do synthetic political narratives appear synchronised across platforms - and if so, could that synchronisation indicate coordinated manipulation? With topic classification across 19 categories and sentiment analysis, the infrastructure can surface patterns relevant to misinformation research.
Public health information monitoring. What proportion of health-related content on social media is synthetic? Does AI-generated health content carry systematically different sentiment profiles than human-authored equivalents - more alarmist, more polarised, more confident? During public health events, how does the synthetic composition of health-related feeds change?
Empirical MIL research. What proportion of a typical user's feed is actually synthetic? How does this vary by platform, by topic category, by time of day, or by week? These are foundational questions for Media and Information Literacy scholarship because the answer affects how users should interpret apparent consensus and credibility in misinformation environments.
Future capability: account-level pattern detection. Currently, Rehumanize operates at the content level - individual posts. However, content-level signals naturally aggregate into account-level patterns. If a high proportion of an account's posts consistently score above the detection threshold, that constitutes a behavioural signal about the account itself - a potential indicator of a fully synthetic or bot-operated account. This is not a current feature, but the data model supports it naturally as the system scales.
Impact and Inclusion
The harm of invisible synthetic content is not evenly distributed. Some populations are systematically more vulnerable to misinformation, and Rehumanize's ambient design is specifically well-suited to expose synthetic voices without requiring advanced media-literacy skills.
People with lower baseline media literacy - those who have not had access to critical thinking training, digital literacy programmes, or higher education - may lack the frameworks to independently evaluate content authenticity. They can be especially vulnerable to synthetic grassroots opinions presented as genuine human testimony. Ambient detection requires no prerequisite knowledge. It works as a visibility layer for both experienced media users and people who have never heard of large language models.
Non-native language speakers face a compounding disadvantage. AI-generated text in a user's second or third language can be harder to evaluate for stylistic tells - subtle phrasing patterns, unnatural fluency, or missing colloquialisms - that native speakers might notice. When synthetic misinformation is embedded in otherwise fluent text, linguistic intuition becomes an unreliable defence. Ambient detection reduces the need for that intuition.
Communities in the Global South often face the weakest platform governance and content moderation while being among the most aggressively targeted by influence operations. Platform investment in content moderation is disproportionately concentrated in English-language markets, leaving non-English-speaking communities with less protection against synthetic content. A user-level detection tool operates independently of platform investment decisions.
Older demographics are less likely to be familiar with AI capabilities and therefore less likely to even suspect that content in their feed might be synthetic. The paste-and-check model of detection is entirely inaccessible to someone who does not know the problem exists. Ambient detection reaches them precisely because it requires no awareness that the problem is there.
Marginalised communities that are already disproportionately targeted by disinformation campaigns stand to benefit from a tool that makes synthetic content visible without requiring the resources, education, or institutional support that are often least available to them. The point is not to decide what is misinformation, but to expose a hidden property of the information environment that can otherwise make misinformation appear more credible or widely supported.
Rehumanize's ambient, passive design is inherently inclusive. It does not require digital sophistication, critical thinking training, institutional access, or even awareness that the problem exists. No curriculum, no workshop, no prerequisite knowledge. That is what makes it MIL infrastructure rather than MIL education - it reaches the people traditional approaches structurally cannot.
Challenges and Risks
The most significant challenge Rehumanize faces is structural, not technical. Ambient passive detection has low direct commercial viability: platforms have little incentive to fund tools that make synthetic content in their feeds visible, while users understandably expect a public-interest transparency tool to be free. However, inference costs, server hosting, and backend infrastructure scale with usage. This is a structural reality of public-interest misinformation infrastructure: the value is social and democratic, while the market incentives are weak. Institutional support - from universities, foundations, and public bodies - may therefore be necessary for sustainability.
A technical constraint is that AI content detection remains probabilistic. The detection model returns a likelihood score, not a binary verdict. Rehumanize's indicators are deliberately framed as "likely AI-generated" rather than definitive determinations. This is a limitation, but it is also a design principle. A probabilistic signal is more intellectually honest than a binary flag - and it is more aligned with MIL principles: users are given information to evaluate critically, including the detection signal itself. In the context of misinformation, the indicator is one piece of evidence about source authenticity, not a verdict about whether a claim is true or false.
This probabilistic limitation is complemented by advances in provider-side watermarking. SynthID-Text, for example, is designed to embed a statistical signal during text generation and provide machine-detectable evidence of AI origin (Dathathri et al., 2024). However, watermarking depends on the generation system applying the mark and can be weakened by changes to the text. Open-source, fine-tuned, or otherwise non-participating systems may not provide a usable watermark. Model-agnostic detection therefore addresses a complementary point in the chain: the user's feed. Together, provider-side marking and user-side detection form layers of a broader transparency response to misinformation and synthetic social influence.
Conclusion
Social media is a major environment through which people understand the world - and that environment is becoming increasingly synthetic. The challenge is no longer limited to the traditional question of whether content is true or false. In the age of AI, misinformation has acquired new dimensions: authenticity of origin, manufactured consensus, and the invisible synthetic population of information environments. The problem is therefore not only false facts; it can also involve false signals about who is speaking and how much genuine human agreement exists.
Many existing responses to misinformation require users to act deliberately - to fact-check, copy and paste, attend a workshop, or exercise vigilance. Rehumanize adds a different layer: it works in the background, already live across Threads, Reddit, and X, quietly making synthetic origin visible while users consume the feed.
Synthetic does not mean false. But invisible synthetic origin means incomplete context. Rehumanize gives users a missing piece of information about the environment they inhabit, allowing them to calibrate their judgement accordingly. That is not a replacement for fact-checking. It is a foundational act of media literacy for an AI era in which misinformation can involve both what a message says and whether the apparent human voice behind it is genuine.
A working prototype is already live. The infrastructure is already built. The goal is simple: when synthetic voices enter the information environment, users should have a way to know.
Reference List
Bai, H., Voelkel, J. G., Muldowney, S., Eichstaedt, J. C., & Willer, R. (2025). LLM-generated messages can persuade humans on policy issues. Nature Communications, 16, 6037. https://doi.org/10.1038/s41467-025-61345-5
Camacho-Collados, J., Rezaee, K., Riahi, T., Ushio, A., Loureiro, D., Antypas, D., Boisson, J., Espinosa-Anke, L., Liu, F., Martínez-Cámara, E., Medina, G., Buhrmann, T., Neves, L., & Barbieri, F. (2022). TweetNLP: Cutting-Edge Natural Language Processing for Social Media. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 38-49. https://doi.org/10.18653/v1/2022.emnlp-demos.5
DataReportal. (2025). Global social media statistics. https://datareportal.com/social-media-users
Antypas, D., Ushio, A., Camacho-Collados, J., Silva, V., Neves, L., & Barbieri, F. (2022). Twitter Topic Classification. Proceedings of the 29th International Conference on Computational Linguistics, 3386-3400. https://aclanthology.org/2022.coling-1.299/
European Commission. (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act). https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
Hackenburg, K., & Margetts, H. (2024). Evaluating the persuasive influence of political microtargeting with large language models. Proceedings of the National Academy of Sciences, 121(24), e2403116121. https://doi.org/10.1073/pnas.2403116121
Jakesch, M., Hancock, J. T., & Naaman, M. (2023). Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences, 120(11), e2208839120. https://doi.org/10.1073/pnas.2208839120
Saab, B. (2024). Manufacturing deceit: How generative AI supercharges information manipulation. National Endowment for Democracy. https://www.ned.org/manufacturing-deceit-how-generative-ai-supercharges-information-manipulation/
Schilke, O., & Reimann, M. (2025). The transparency dilemma: How AI disclosure erodes trust. Organizational Behavior and Human Decision Processes, 188, 104405. https://doi.org/10.1016/j.obhdp.2025.104405
Dathathri, S., See, A., Ghaisas, S., Huang, P.-S., McAdam, R., Welbl, J., et al. (2024). Scalable watermarking for identifying large language model outputs. Nature, 634, 818-823. https://doi.org/10.1038/s41586-024-08025-4
Sun, Z., Zhang, Z., Shen, X., Zhang, Z., Liu, Y., Backes, M., Zhang, Y., & He, X. (2025). Are we in the AI-generated text world already? Quantifying and monitoring AIGT on social media. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 22975-23005. https://aclanthology.org/2025.acl-long.1120/
Udry, J., & Barber, S. J. (2024). The illusory truth effect: A review of how repetition increases belief in misinformation. Current Opinion in Psychology, 56, 101736. https://doi.org/10.1016/j.copsyc.2023.101736
Wack, M., Ehrett, C., Linvill, D., & Warren, P. (2025). Generative propaganda: Evidence of AI's impact from a state-backed disinformation campaign. PNAS Nexus, 4(4), pgaf083. https://doi.org/10.1093/pnasnexus/pgaf083
Woolley, S. (2023). Manufacturing consensus: Understanding propaganda in the era of automation and anonymity. Yale University Press. https://doi.org/10.12987/yale/9780300251234.001.0001