Guardrail Vulnerabilities in Open-Source Language Models: Implications for Democratic Discourse and Marginalized Communities
| dc.contributor.author | Münker, Simon | |
| dc.contributor.author | Sartori, Fabio | |
| dc.date.accessioned | 2025-12-23T16:39:58Z | |
| dc.date.available | 2025-12-23T16:39:58Z | |
| dc.date.issued | 2026-01-06 | |
| dc.description.abstract | The proliferation of open-source Large Language Models (LLMs) presents a complex technological phenomenon with significant societal implications. While these models democratize access to advanced Natural Language Processing (NLP) capabilities, they simultaneously amplify risks for marginalized communities who often bear the disproportionate burden of technological misuse. Our research examines systematic vulnerabilities in guardrail mechanisms across seven prominent open-source LLMs, revealing patterns of harmful content generation that threaten democratic discourse and social cohesion. Through empirical analysis using advanced NLP classification methods, we demonstrate that popular open-source models consistently generate content classified as hateful or offensive when subjected to adversarial prompting techniques. These findings directly contradict the safety assurances provided by model developers, particularly Meta AI's stated commitment that their systems should present balanced perspectives on debated policy issues rather than singular viewpoints. | |
| dc.format.extent | 10 pages | |
| dc.identifier.doi | https://doi.org/10.24251/HICSS.2026.804 | |
| dc.identifier.isbn | 978-0-9981331-9-5 | |
| dc.identifier.other | ba46f294-65e8-4c3b-8580-2761042d0ba1 | |
| dc.identifier.uri | https://hdl.handle.net/10125/112207 | |
| dc.language.iso | eng | |
| dc.relation.ispartof | Proceedings of the 59th Hawaii International Conference on System Sciences | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | |
| dc.rights.uri | https://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Digital Democracy and Social Cohesion | |
| dc.subject | ai safety | |
| dc.subject | democratic discourse | |
| dc.subject | guardrail vulnerabilities | |
| dc.subject | hate speech detection | |
| dc.subject | language models | |
| dc.title | Guardrail Vulnerabilities in Open-Source Language Models: Implications for Democratic Discourse and Marginalized Communities | |
| dc.type | Conference Paper | |
| dc.type.dcmi | Text | |
| prism.startingpage | 6792 |
Files
Original bundle
1 - 1 of 1
