Guardrail Vulnerabilities in Open-Source Language Models: Implications for Democratic Discourse and Marginalized Communities

dc.contributor.authorMünker, Simon
dc.contributor.authorSartori, Fabio
dc.date.accessioned2025-12-23T16:39:58Z
dc.date.available2025-12-23T16:39:58Z
dc.date.issued2026-01-06
dc.description.abstractThe proliferation of open-source Large Language Models (LLMs) presents a complex technological phenomenon with significant societal implications. While these models democratize access to advanced Natural Language Processing (NLP) capabilities, they simultaneously amplify risks for marginalized communities who often bear the disproportionate burden of technological misuse. Our research examines systematic vulnerabilities in guardrail mechanisms across seven prominent open-source LLMs, revealing patterns of harmful content generation that threaten democratic discourse and social cohesion. Through empirical analysis using advanced NLP classification methods, we demonstrate that popular open-source models consistently generate content classified as hateful or offensive when subjected to adversarial prompting techniques. These findings directly contradict the safety assurances provided by model developers, particularly Meta AI's stated commitment that their systems should present balanced perspectives on debated policy issues rather than singular viewpoints.
dc.format.extent10 pages
dc.identifier.doihttps://doi.org/10.24251/HICSS.2026.804
dc.identifier.isbn978-0-9981331-9-5
dc.identifier.otherba46f294-65e8-4c3b-8580-2761042d0ba1
dc.identifier.urihttps://hdl.handle.net/10125/112207
dc.language.isoeng
dc.relation.ispartofProceedings of the 59th Hawaii International Conference on System Sciences
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 International
dc.rights.urihttps://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectDigital Democracy and Social Cohesion
dc.subjectai safety
dc.subjectdemocratic discourse
dc.subjectguardrail vulnerabilities
dc.subjecthate speech detection
dc.subjectlanguage models
dc.titleGuardrail Vulnerabilities in Open-Source Language Models: Implications for Democratic Discourse and Marginalized Communities
dc.typeConference Paper
dc.type.dcmiText
prism.startingpage6792

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
0663.pdf
Size:
200.16 KB
Format:
Adobe Portable Document Format