An overview of our work on privacy-aware NLP: federated learning for Algerian social media sentiment, comparing FedAvg and FedProx under highly non-IID client conditions.
All links open in a new tab.
Why applying federated learning to Algerian dialect NLP is both challenging and worthwhile.
How the data were collected, split, and distributed across simulated federated clients.
FedAvg is the most interpretable starting point for any federated learning benchmark.
FedProx was selected to address the update drift risk that arises from severely imbalanced client data.
The chosen values prioritise a credible first benchmark rather than aggressive tuning.
These are reasonable benchmark choices, not fully optimised hyperparameters. The paper explicitly identifies μ sweeps, α sweeps, and multi-seed evaluation as future work. The contribution is a strong first benchmark, not a final tuned system.
FedProx narrows the gap to centralized training to under one percentage point.