MSc research · Multi-agent · Voice
HireAgent: Multilingual Multi-Agent Job Platform
MSc Data Science final research · Coventry University UK (NIBM) · 2025/26
A conversational AI platform for Sri Lanka's informal employers. They speak or text a job in any local language, and a multi-agent pipeline turns it into a moderated, published listing. No forms, and no English required.
- Role
- Researcher · AI Engineer
- Year
- 2025 – 26
- Focus
- MSc research
- Stack
- 8 technologies
- CrewAI
- OpenAI
- MCP
- FastAPI
- Twilio Voice
- pgvector
- Redis
- Railway
0.83
Extraction F1 (P 0.87 / R 0.80)
433ms
Real-time search latency
1.00
Scam-detection recall
92.5%
Intent routing accuracy
95.8%
Functional tests passed (23/24)
0
False positives on clean inputs
The problem
About 70% of employment in Sri Lanka is informal. Employers hire by word of mouth, signage, and WhatsApp, while job portals are desktop-first, English-only, and full of forms. Mobile internet reaches 80% of people, but computer literacy is only 36.4% and desktop ownership 19.5%. On top of that, unmoderated job posts expose job seekers to fraud and exploitation.
Overview
I designed, built, and evaluated a multi-agent platform to remove those barriers. It accepts voice calls, SMS, WhatsApp, and web chat. It detects language and intent, asks clarifying questions, and pulls out a structured job posting. A six-stage CrewAI pipeline then validates, enriches, classifies, moderates, and publishes the posting.
The research followed Design Science Research (Hevner et al., 2004) with iterative prototyping: 8 sprints over 14 weeks. I tested it on scenario corpora covering complete English, fragmented Singlish, Tamil-English mixes, Romanized Sinhala, and deliberate scam and discriminatory inputs.
HireAgent began as a single-repository prototype. It grew into a distributed system of four core repositories (published as BrokeMe) deployed on Railway as 7 services.
Architecture
End-to-end architecture
Real-time and asynchronous work are split. Searches return in under 500ms, while AI listing creation (about 7.6s) runs in the background so the user never waits.
- 01
Omnichannel intake
Voice · SMS · WhatsApp · WebTwilio ConversationRelay for live calls, plus SMS, WhatsApp, and web chat.
- 02
Understanding layer
brokeme-intentLanguage detection, intent detection, and conversational extraction into a structured schema, with session state in Redis.
- 03
Bifurcated routing
433ms vs asyncCustomer searches take the real-time path. New listings go onto a Redis queue.
- 04
Worker
brokeme-workerThree daemon threads consume the queue and hand jobs to the agent pipeline.
- 05
Multi-agent pipeline
brokeme-agentsSix CrewAI stages with defensive JSON parsing, so hallucinated output can't corrupt later stages.
- 06
Publish & notify
brokeme-apiFastAPI and PostgreSQL with pgvector store the listing, publish it, and notify the user.
How it works
Six-stage agent pipeline
Each stage has one job and a tuned temperature: low for checks, higher for creative writing.
- 01
Validation
Stage 1 · temp 0.1Checks the posting is complete across all 8 schema fields.
- 02
Enrichment
Stage 2 · temp 0.3Adds tags, synonyms, and SEO keywords.
- 03
Classification
Stage 3 · temp 0.2Assigns one of 45 categories.
- 04
Moderation
Stage 4 · temp 0.1Screens for scams and discriminatory language before anything is published.
- 05
Content generation
Stage 5 · temp 0.4Writes professional, ready-to-publish ad copy.
- 06
Image prompt
Stage 6 · temp 0.4Generates a DALL·E prompt for the listing visual.
Components
brokeme-intent
Channel gateway: voice, SMS, WhatsApp, and web intake, language and intent detection, structured extraction.
brokeme-api
FastAPI with clean architecture, PostgreSQL with pgvector, OTP auth, and service and listing management.
brokeme-worker
Consumes the Redis queue and runs AI processing off the request path.
brokeme-agents
The CrewAI six-stage pipeline: validation, enrichment, classification, moderation, content, and image prompt.
provider-search-mcp
An MCP server that gives agents safe, read-only tools: search providers, get details, check slots.
Results
Extraction F1 by language
Higher is better. Code-mixed and romanized input is the hardest.
- Standard English0.93
- Partial English0.83
- Tamil-English0.75
- Code-mix Singlish0.74
- Romanized Sinhala0.66
Speech recognition word error rate
Lower is better. Controlled 16kHz audio, 19.1% overall.
- Clear English7.2%
- Sinhala21.4%
- Singlish mix28.6%
Research contributions
- 01Evidence that a multi-agent conversational design lowers cognitive barriers for low-literacy informal employers
- 02Empirical speech recognition benchmarks for Sinhala and code-mixed Singlish under controlled conditions
- 03An architecture pattern that splits real-time search from asynchronous AI processing
- 04An inclusive AI design framework for multilingual, mobile-first developing economies
Limitations & future work
- Voice was tested through WebSocket simulation, not a live mobile network
- Agent tests used mocks, so they validate the architecture rather than the meaning of outputs
- Recall on discriminatory language is 0.80, because implicit bias is harder to catch
- Next: a formal usability study with Sri Lankan SME owners, token streaming for sub-200ms voice replies, and speech recognition fine-tuned on local dialects