Lecture materials for CSC 84040 will be posted here as they become available. Students should check this page regularly for updates and supporting materials related to distributed mining, agentic AI, and modern scientific-discovery workflows.

  • Introduction to Data Mining at Scale and AI Agents
    tl;dr: Distributed data-mining foundations and large-scale systems for AI agents.
    [slides]

    Introduction to the course, its goals, and the foundational concepts of distributed systems and data mining. This lecture sets the stage for understanding how large-scale data processing frameworks can be used in conjunction with AI agents to perform complex tasks efficiently.

  • Frequent Itemsets & Data Quality: Apriori, FP-Growth, and Identifying Statistical Properties for Usable Data
    tl;dr: Mining frequent patterns while evaluating whether data is actually suitable for downstream learning.

    This class covers frequent-itemset mining, pattern extraction, and the statistical properties that determine whether data is reliable and useful in practice.

  • Similarity Search & LSH: Semantic Vector Search & Agent Memory Deduplication
    tl;dr: Approximate nearest-neighbor search and memory deduplication in agentic systems.

    Students examine locality-sensitive hashing, semantic search, and deduplication strategies for agent memory and retrieval pipelines.

  • Data Stream Mining & System Monitoring: Sensor Stream Mining, Energy Signals, and Memory Parameters (DGIM, Sketches)
    tl;dr: Streaming algorithms for monitoring sensor and system data under tight memory constraints.

    We focus on DGIM, sketches, and stream-monitoring concepts for robotic, sensor, and energy-aware systems operating in real time.

  • Link Analysis & Web Mining: Citation Network Mining & SciSciNet (PageRank, HITS)
    tl;dr: Graph-based mining and citation analysis for literature and web-scale networks.

    This lecture introduces PageRank, HITS, and citation-network mining as tools for understanding complex research and information networks.

  • Clustering for Massive Data & Vision Problems: Hypothesis Clustering, Proximity Agents, and Vision-Based Open Problems
    tl;dr: Large-scale clustering and visual understanding in data-mining and agentic research workflows.

    Students explore clustering for large data and the role of proximity-based methods in vision and multimodal problem settings.

  • Game Theory & Equilibrium in Data Mining: Coordination, Competition, and Practical Control Capabilities
    tl;dr: How strategic interactions and equilibrium concepts inform multi-agent and control-oriented data systems.

    This lecture connects game theory, coordination, and practical control to the design of robust data-driven multi-agent systems.

  • Evolutionary Learning, Mean-Field Learning, Self-play for Self-Improvement and Continual Learning
    tl;dr: Adaptive learning paradigms for long-lived, self-improving agent systems.

    The lecture covers evolutionary and self-play learning approaches to continual adaptation in real-time environments.

  • Classification, Prediction, & Dimensionality: Concept Embedding & Subspace Discovery
    tl;dr: Prediction and dimensionality reduction as tools for concept discovery and contextual modeling.

    This session explores classification, prediction, embedding representations, and subspace discovery for high-dimensional data.

  • Graph Mining & Network Analysis: Large-Scale GNNs & SNAP Benchmarks
    tl;dr: Graph learning, large-scale network mining, and evaluation against benchmark datasets.

    Students investigate graph mining, network analysis, and the role of GNNs and graph benchmarks in modern data-mining workflows.

  • Data Optimization: Determining Data Requirements for Specific Optimization Problems & TimesFM
    tl;dr: Understanding when and how much data is sufficient for optimization and forecasting tasks.

    This lecture addresses data requirements, optimization-based decision making, and forecasting models such as TimesFM.

  • Item Response Theory & LBD: Ranking Text-Based Data, Swanson's ABC Model & Hypothesis Generation
    tl;dr: Ranking data quality and discovering scientific hypotheses through literature-based discovery.

    The class covers item response theory, text ranking, and literature-based discovery methods for hypothesis generation.

  • Multi-Agent Scientific Discovery: Architecture of AI Co-Scientist Systems and Elo-based Tournament Evolution
    tl;dr: How multi-agent systems can propose, test, and refine hypotheses in scientific workflows.

    This session examines AI co-scientist architectures, agent debate, and evaluation frameworks such as Elo-based tournament evolution.

  • Knowledge Graphs & GraphRAG: Hierarchical Community Detection, ULTRA, & Inference
    tl;dr: Knowledge graphs, graph retrieval, and hierarchical community-aware inference.

    We examine knowledge-graph reasoning, graph retrieval augmentation, and hierarchical community detection for robust inference.

  • Reliability, Ethics, & Evaluation: Epistemic Honesty & Overcoming Dated Data Practices
    tl;dr: Critical evaluation of AI-enabled mining, reliability, and responsible research practice.

    This session focuses on epistemic honesty, replication concerns, ghost evidence, and rigorous evaluation in agent-driven research systems.

  • Advanced Topics: Human-Agent and Agent-Agent Coordination and Competition
    tl;dr: Human-in-the-loop and multi-agent coordination dynamics in decision systems.

    The final lecture explores coordination, competition, and control patterns across human-agent and agent-agent systems.

  • Final Project Presentations: Comprehensive Application of Course Concepts
    tl;dr: Student project presentations that synthesize the semester’s methods and findings.

    Students present their final projects, showing how they applied data-mining, agentic design, evaluation, and research reasoning throughout the semester.