Get in Touch

Course Outline

Day 1: Build the Foundation — Ingest, Search, Retrieve

Module 1: The Legal Engineer’s Landscape

  • Learning objectives—understand the role, AI’s position in legal work, and two pervasive risks.
  • Topics:
    • The legal-engineer role and current market demand.
    • AI applications: eDiscovery, review, contracts, research, investigations; understanding the EDRM model clearly.
    • Build vs. buy considerations.
    • Two universal risks: confidentiality/privilege and defensibility.

Module 2: Legal Data Is Messy — Ingestion and Extraction

  • Learning objectives—manage legal data realities at scale.
  • Topics:
    • Handling 1,400+ file types, email/PST archives, scanned paper, load files (.dat/.opt); understanding critical embedded metadata.
    • Text extraction (Tika), OCR, and deduplication strategies.
  • Lab: FreeEed Ingestion — build an ingestion pipeline over a deliberately messy document set (email/PST, scans, load files).

Module 3: Search and Retrieval — the Foundation

  • Learning objectives—construct the core eDiscovery primitive: find anything within everything.
  • Topics—full-text search and indexing (Solr/Lucene); relevance, metadata, and date filtering; searching across OCR-processed content.
  • Lab: eDiscovery Search — index a corpus and execute real eDiscovery-style searches, including within OCR’d scans.

Module 4: RAG for Legal Documents — with Citations

  • Learning objectives—build RAG over legal documents that cites sources.
  • Topics:
    • Why retrieval, not fine-tuning, is preferred for sensitive material—the model never absorbs the documents directly.
    • Chunking, embeddings, and crucially, citations/provenance.
    • Multi-document and thread summarization.
  • Lab: Legal RAG with Citations — build a RAG Q&A system over a document set that answers with source citations.

Day 2: Make It Private, Defensible, and Shippable

Module 5: Privacy, Privilege, and Local Serving — the Privilege Trap

  • Learning objectives—keep legal data local and certifiable.
  • Topics:
    • Data flow when interacting with cloud AI services.
    • Privilege waiver, duty of competence, and the 'private' spectrum (contractual vs. physical).
    • Morgan v. V2X precedent and why local deployment is court-defensible.
    • Serving local models (Ollama/vLLM) and monitoring outbound traffic.
  • Lab: Local Model + Egress Proof — run a local model end-to-end and prove via monitoring that no data exited the environment.

Module 6: Defensible AI Review

  • Learning objectives—measure and document an AI review to ensure it withstands legal challenge.
  • Topics:
    • Court-admissible metrics: recall, elusion, precision, ground-truth validation; TAR/active learning.
    • Transparency (why was this document coded this way?) and reproducibility—pin the model, fix settings, log everything.
    • The 'defensible case snapshot' allowing a review to be re-run later with identical results.
  • Lab: Defensible Review — measure an AI review against a blind ground truth and produce a reproducibility bundle.

Module 7: Ship It — Workflow, Private Deployment, and Governance

  • Learning objectives—assemble components into a workflow, deploy privately, and score the system.
  • Topics:
    • A multi-step legal workflow (ingest → search → summarize → review → produce) with human-in-the-loop.
    • Private/on-prem deployment essentials (containerization; keeping data on-site).
    • AI governance for legal professionals and scoring the system using SAIS-100 (the Elephant Scale Secure AI Score).
  • Lab: Score and Package — wire a multi-step workflow, score it with SAIS-100, and package it for private deployment.

Capstone (integrated across Day 2)

  • Build a private, defensible legal-AI application end-to-end—ingest a messy corpus, search it, answer questions using citations via a local model, measure defensible review metrics, and package for private deployment.
  • Participants leave with a portfolio project that mirrors actual legal-engineer responsibilities.

Optional Day 3 / Advanced Modules (deliverable as a 3rd day or a modular series)

  • Investigations: Entities, Relationships, and Timelines — extract people/organizations/dates, reconstruct email threads, build chronologies, map near-duplicates and document lineage. Lab: build a timeline and entity/relationship view.
  • Agentic and Multi-Step Legal Workflows (deep dive) — advanced orchestration, contract analysis, multi-document synthesis, tool use, and guardrails as design principles. Lab: build a multi-step workflow with a human checkpoint.
  • Deployment at Scale — on-premises and appliance deployment, distributed processing for large volumes, regulated environments (CJIS, government, higher education), hardware sizing. Lab: containerize and scale a processing job across workers.
  • Governance and Compliance Deep-Dive — AI regulation landscape (100+ US state AI laws, EU AI Act), audit requirements, and a comprehensive SAIS-100 governance audit. Lab: audit a legal-AI system against a governance/defensibility checklist.

Requirements

  • Proficiency in Python and basic APIs.
  • Helpful: User-level familiarity with Large Language Models (LLMs)—no ML background is required as we build the conceptual framework.
  • No legal background required—necessary legal concepts are taught within context.

Audience

  • Software and AI engineers transitioning into legal technology.
  • Engineers at legal-tech companies needing deeper domain knowledge.
  • Technically-minded legal, eDiscovery, or information governance professionals who prefer building over buying.
  • Anyone aiming for the 'legal engineer' or 'AI legal engineer' role.
 14 Hours

Testimonials (1)

Related Categories