Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Day 1: Building the Foundation — Ingest, Search, Retrieve
Module 1: The Legal Engineer’s Landscape
- Learning objectives—understand the role, where AI fits in legal work, and the two critical risks that permeate everything.
- Topics
-
- The legal-engineer role and current market demand
- Where AI fits: eDiscovery, review, contracts, research, investigations; the EDRM model explained simply
- Build vs. buy considerations
- The two core risks: confidentiality/privilege and defensibility
Module 2: Legal Data Is Messy — Ingestion and Extraction
- Learning objectives—handle the reality of legal data at scale.
- Topics
- 1,400+ file types, email and PST formats, scanned paper, load files (.dat/.opt); critical embedded metadata
- Text extraction (Tika), OCR, and deduplication as strategic choices
- Lab: FreeEed Ingestion—build an ingestion pipeline over a deliberately messy document set (email/PST, scans, load files)
Module 3: Search and Retrieval — The Foundation
- Learning objectives—construct the core eDiscovery primitive: finding anything inside everything.
- Topics—full-text search and indexing (Solr/Lucene); relevance, metadata and date filtering; searching across OCR’d content
- Lab: eDiscovery Search—index a corpus and run real eDiscovery-style searches, including within OCR’d scans
Module 4: RAG for Legal Documents — With Citations
- Learning objectives—build RAG over legal documents that cites its sources.
- Topics
- Why retrieval, not fine-tuning, is preferred for sensitive material—the model never consumes the documents directly
- Chunking, embeddings, and crucially citations/provenance
- Multi-document and thread summarization
- Lab: Legal RAG with Citations—build a RAG Q&A over a document set that answers with source citations
Day 2: Making It Private, Defensible, and Shippable
Module 5: Privacy, Privilege, and Local Serving — The Privilege Trap
- Learning objectives—keep legal data local and certifiable.
- Topics
- Where data goes when it reaches a cloud AI
- Privilege waiver, duty of competence, and the “private” spectrum (contractual vs. physical)
- Morgan v. V2X case study and why local deployment is court-defensible
- Serving local models (Ollama/vLLM) and monitoring outbound traffic
- Lab: Local Model + Egress Proof—run a local model end-to-end and prove, via monitoring, that no data left the premises
Module 6: Defensible AI Review
- Learning objectives—measure and document an AI review so it withstands legal scrutiny.
- Topics
- Court-worthy metrics: recall, elusion, precision, ground-truth validation; TAR/active learning
- Transparency (why did it code this document?) and reproducibility—pin the model, fix settings, log everything
- The “defensible case snapshot” enabling someone to re-run your review a year later with identical results
- Lab: Defensible Review—measure an AI review against a blind ground truth and produce a reproducibility bundle
Module 7: Ship It — Workflow, Private Deployment, and Governance
- Learning objectives—assemble components into a workflow, deploy privately, and score the system.
- Topics
- A multi-step legal workflow (ingest → search → summarize → review → produce) with human-in-the-loop
- Private/on-premises deployment essentials (containerization; keeping data in-house)
- AI governance for legal in brief, and scoring the system using SAIS-100 (the Elephant Scale Secure AI Score)
- Lab: Score and Package—wire a multi-step workflow, score it with SAIS-100, and package it for private deployment
Capstone (integrated across Day 2)
- Build a private, defensible legal-AI application end to end—ingest a messy corpus, search it, answer questions with citations using a local model, measure a defensible review, and package for private deployment.
- Participants leave with a portfolio project that mirrors the legal-engineer job role.
Optional Day 3 / Advanced Modules (deliverable as a 3rd day or a modular series)
- Investigations: Entities, Relationships, and Timelines—extract people/orgs/dates, reconstruct email threads, build chronologies, map near-duplicates and document lineage. Lab: build a timeline and entity/relationship view.
- Agentic and Multi-Step Legal Workflows (deep)—richer orchestration, contract analysis, multi-doc synthesis, tool use, and guardrails as a design principle. Lab: build a multi-step workflow with a human checkpoint.
- Deployment at Scale—on-premises and appliance deployment, distributed processing for large volumes, regulated environments (CJIS, government, higher-ed), hardware sizing. Lab: containerize and scale a processing job across workers.
- Governance and Compliance Deep-Dive—the AI-regulation landscape (100+ US state AI laws, the EU AI Act), audit requirements, and a full SAIS-100 governance audit. Lab: audit a legal-AI system against a governance/defensibility checklist.
Requirements
- Familiarity with Python and basic APIs
- Helpful but not required: User-level familiarity with LLMs (no ML background needed—we build the mental model)
- No legal background required—necessary legal concepts are taught in context
Audience
- Software and AI engineers transitioning into legal technology
- Engineers at legal-tech companies needing deeper domain expertise
- Technically-oriented legal, eDiscovery, or information-governance professionals who want to build rather than just buy
- Anyone aiming for the “legal engineer” or “AI legal engineer” role
14 Hours
Testimonials (2)
The session was highly interactive and applicable to the business.
Jorge Boscan - Chevron Global Technology Services Company
Course - Advanced GitHub Copilot & AI for Projects and Infrastructure
Machine Translated
That i gained a knowledge regarding streamlit library from python and for sure i'll try to use it to improve applications in my team which are made in R shiny