AI 401 · Year 4 · Semester 1 · 3 credits · AI & Machine Learning
NLP & Large Language Models for Insurance
Transforming unstructured insurance text into auditable actuarial intelligence—from policy wordings to claims triage.
Latent Coverage: The NLP Reserving Engine
When catastrophic storm backlogs and unstructured claims notes threaten a commercial insurer with insolvency and regulatory sanctions, lead actuarial data scientist Maya Lin must engineer an end-to-end, mathematically governed LLM pipeline to triage fifty thousand files before the quarterly board audit.
Maya arrives at her desk at six in the morning to find the claims floor in crisis. Forty-two thousand first notice of loss notes from a catastrophic midwest hail and wind storm are piled up as free text, and the legacy bag-of-words model just classified three thousand severe commercial roof collapses as routine glass claims because of misspellings and out-of-vocabulary jargon.
Transcript
Maya arrives at her desk at six in the morning to find the claims floor in crisis. Forty-two thousand first notice of loss notes from a catastrophic midwest hail and wind storm are piled up as free text, and the legacy bag-of-words model just classified three thousand severe commercial roof collapses as routine glass claims because of misspellings and out-of-vocabulary jargon.
- Contrast rule-based, word-level, and subword (Byte Pair Encoding) tokenisation schemes on unstructured insurance text.
- Derive the embedding lookup mechanism and formulate the Skip-Gram with Negative Sampling objective.
- Compute cosine similarity and document-level pooled embeddings to quantify semantic distance between adjuster notes.
- Evaluate embedding representations under Actuarial Standards of Practice regarding data quality and model governance.