README.md
Mohannad Ibrahim
I build AI agents and the backends that keep them reliable.
- $650,000new annualized premium attributed to an agency-retention modelproof ↓ c4
- 28%less manual call handling, via a voice agent on MCP toolsproof ↓ c2
- p95 <200 msacross ~400,000 monthly requests on an eligibility APIproof ↓ c3
Case studies
Five builds with numbers attached, grouped by what they are. Open the details here, or read any of them in the console as c1 to c5.
Agents 02
-
Conagra Brands2025-26
Production agent with Graph RAG
ProductionA Copilot Studio assistant in Microsoft Teams that hands multi-step reasoning to a custom LangGraph backend, with Graph RAG on Neo4j.
Est. 10-15 staff hours saved per week/Copilot Studio evaluation matrix
The assistant lives in Microsoft Teams. Copilot Studio hands multi-step reasoning to a custom LangGraph Python backend over Microsoft Entra-secured APIs.
Standard vector search fell short on deeply nested SharePoint documents, so Graph RAG on Neo4j lets it traverse the document hierarchy and map supply-chain entity relationships. Tested with a Copilot Studio evaluation matrix.
Software Engineer, Raikes Design Studio capstone. Aug 2025 - May 2026.
-
Ameritas2025-26
Voice agent + MCP tool gateway
ShippedA Python/Flask backend for a VAPI voice agent that handles monthly password resets and account lookups.
28% less manual call handling/MCP over JSON-RPC 2.0/PII stays server-side
The agent automates monthly password resets and account lookups against Active Directory and the Mainframe.
Those actions are exposed as MCP tools over JSON-RPC 2.0 with least privilege, so PII stays server-side. Manual call handling dropped 28%.
Backend 01
-
Ameritas2025-26
Eligibility API reliability
ProductionOwned a C#/.NET eligibility-and-claim-status service and made it fast, observable, and releasable on demand.
p95 <200 ms/~400,000 requests a month/test coverage ~40% to 85%
Moved the service to ECS Fargate, which took releases from a monthly window to on demand. CloudWatch metrics and alarms cut time to detect a class of eligibility mismatch from days to under 15 minutes.
Unit-test coverage went from ~40% to 85%. After a senior engineer flagged stale-eligibility risk, I revised the caching design to a short-TTL cache keyed on member.
Data 02
-
Ameritas2025-26
Agency-retention model
PilotedA model that flags flight-risk agencies about 6 months early, delivered as a Power BI "Hit List".
$650,000 new annualized premium attributed/0.81 AUC
Built a 22,000-record training set (~1,500 agencies, 4 states) from public DOI filings, NIPR, and M&A releases. LightGBM via PyCaret reached 0.81 AUC.
A Power BI "Hit List" flags flight-risk agencies ~6 months early. In a 3-month regional pilot, 3 high-producing agencies onboarded: $650,000 new annualized premium attributed.
-
Personalopen repo
AEC Opportunity Pipeline
Open repoFour live public sources conformed into one Snowflake medallion star schema.
358 pytest tests/1,957 records to 1,952 opportunities, 0 false merges/55 checks per load
Sources: SAM.gov API, Nebraska DOT PDF, City of San Diego CIP CSV, and Texas DOT Socrata API.
Entity resolution takes 1,957 records into 1,952 opportunities with false merges at zero. 55 data-quality checks are logged per load, and 358 pytest tests cover it.
More work
Smaller pieces, same habits: make failure visible, then make it rare.
-
Ingestion at ~2M events a day Conagra
Idempotency keys on DynamoDB writes, a dead-letter queue, and CloudWatch custom metrics. Diagnosing lag went from a manual log grep to a five-minute dashboard check.
-
One warehouse for 14 source feeds UNL ADMA
14 source feeds into one Aurora PostgreSQL warehouse (S3, Glue). The reporting cycle went from two days to under one hour, and data incidents from roughly weekly to near zero. LangSmith tracing of multi-agent runs; MLflow Tracking, Projects, Models, and Registry.
-
Meeting Notes RAG Assistant Personal
Tool-calling RAG over meeting transcripts, with citations to meeting, speaker, and timestamp. It returns "not found" instead of guessing.
Illustrative run, shows behavior, not real logs- plan"What did we decide about the vendor contract?" Search meetings, then answer with citations.
- tool callsearch_meetings(query, date range)
- 0 resultsNothing in that range. Widen it instead of guessing.
- retrysearch_meetings(query, broader date range)
- 2 hitsTwo meetings mention the vendor contract.
- answerAnswers with citations: meeting, speaker, timestamp. If nothing matched, it says "not found".
Work
- Jul 2026 - PresentHandshake AI FellowHandshake AI
Evals for top AI labs: I write eval tasks, rubrics, and graders, grade model outputs, and red-team models.
- Feb 2026 - PresentData EngineerUniversity of Nebraska-Lincoln, ADMA
- Nov 2025 - May 2026Data and AI Engineer InternAmeritas
- Aug 2025 - May 2026Software Engineer, Raikes Design Studio capstoneConagra Brands
- May 2025 - Nov 2025Software Developer InternAmeritas
Education
B.S. in Computer Science
University of Nebraska-Lincoln, May 2026
Associate of the Jeffrey S. Raikes School of Computer Science and Management
GPA 3.9
"One of the most exceptional interns I have ever worked with."
Contact
Lincoln, NE. Open to Bay Area roles. Available now.
mibrahimhg@gmail.com