🌐 Founder Website (Almadi Sahara) 💼 LSI LinkedIn
Open-Source AI Governance & Multilingual Trust

Linguistic Security Institute

The Linguistic Security Institute (LSI) is an open-source research initiative dedicated to ensuring AI systems are culturally coherent, institutionally resilient, and ethically verifiable. We audit linguistic bias in AI, develop open-source tools for bias detection, create educational resources for developers and students, and propose governance frameworks for linguistic safety in AI development.

Core Focus Pillars

Bias Auditing & Detection

Rigorous evaluation of LLMs for dialectal degradation, tokenization disparities, and semantic failure modes across low-resource and high-resource non-English languages.

DialectologyRobustness

Open-Source Tooling

Building accessible Python toolkits, security frameworks, and datasets designed to help developers shield applications from polyglot poisoning and cross-lingual prompt injections.

PythonFirewalls

Governance & Education

Proposing standards for ethical linguistic safety in institutional AI pipelines and running hands-on security workshops for students and researchers.

PolicyWorkshops

Active Initiatives

Linguistic Firewall: Polyglot Poisoning

An open-source Python toolkit engineered for detecting multilingual prompt injection, safety filter bypasses via code-switching (Spanglish, Arabizi), and structural polyglot vulnerabilities in frontier models.

SecurityMultilingual NLP

Mitote

A specialized RAG text-to-speech prototype focused on Indigenous language preservation, supporting Nahuatl-aware Spanish pronunciation mappings and regional prosody features.

Audio AIPreservation

Selected Publications & Benchmarks

BAREC Shared Task 2026

LSI at BAREC Shared Task 2026: A Controlled Study of Training Convergence, Heuristics, and Ensembles for Arabic Sentence-Level Readability

Ranked 1st in the Constrained track and 3rd in the Open track with an ensemble score using logit-level ensembles and Earth Mover's Distance loss.

View BAREC Shared Task →
SemEval-2026 Task 9 (POLAR)

NAMAA at SemEval-2026 Task 9: Comparing Generative, Retrieval-Augmented, and Discriminative Methods for Arabic Online Polarization Detection and Type Classification

Advanced architectures for multi-label type classification combining encoder fine-tuning, zero-shot prompting, and retrieval-augmented in-context learning (RAG-ICL).

View ACL Anthology Paper →

Upcoming Workshops & Demos

Official Hands-On Workshop

Linguistic Firewall: Expanded Reproducible Version

An expanded, fully reproducible workshop built on top of the security demo presented in April 2026. Join us for an intensive hands-on session covering polyglot red-teaming and defense frameworks.

📍 DEATHCon BSides — November 14, 2026

View DeathCon Workshops →