Project Overview & Accomplishments


This ARO MURI project has successfully concluded. It advanced the foundations for robust and adaptive machine learning inspired by childhood development, addressing vulnerabilities to adversarial inputs (malicious and non-malicious) at both design-time and runtime. The work tackled key challenges: lack of robustness to sensor noise, image transformations, and distribution shifts; the need for concepts built on prior knowledge; and verification/monitoring in dynamic, safety-critical environments.

Our DARPA L2M-inspired evaluations used a custom robotic platform as a battlefield surrogate, with tasks framed as adversarial scavenger hunts integrating navigation, perception, and control. ICRA 2022 live demonstrations showcased real-time lifelong learning under disruptions, using ensembling and replay regularization for robust occupancy prediction despite occlusions—these scaled to comprehensive evaluations across all thrusts.


Thrust I: Concept-based Learning Robust to Adversarial Examples

This thrust improved deep learning robustness through child-like reasoning in three complementary approaches, yielding state-of-the-art results on benchmarks and robotic platforms.

  • First Approach: Robust Perception-Based Recognition Invariant to Geometric Transformations
    Developed the Transformed Risk Minimization (TRM) framework, with a zero-gap theorem and PAC-Bayes bounds linking transformed risks to learning theory. The SCALE algorithm automatically learned task-adaptive augmentations (e.g., rotations, flips) for continuous/discrete transforms. Extensions supported uncertainty estimation and adversarial detection. An SE(3)-equivariant network with attention-based local shape modeling outperformed SO(3)-equivariant and non-equivariant baselines on large-scale, unoriented, noisy point clouds without pre-segmentation. Symmetry learning algorithms surpassed handcrafted methods, and conditional 3D pose distribution networks improved downstream tasks under ambiguity. MOLAR advanced collaborative regression/bandits with tighter error/regret bounds. Evaluated on CIFAR-10 and robots, these ensured strong invariance to geometric adversaries.
  • Second Approach: Inductive Biases for Robustness to Dimensional Alterations
    Engineered representations prioritizing shapes and relationships (e.g., black/white encodings, video background subtraction), effective against stochastic and adversarial attacks. Adversarial training synthesized robust symbols. Imprecise Bayesian Neural Networks (IBNNs) with credal sets outperformed standard BNNs on CIFAR-10, MNIST, Fashion-MNIST, SVHN under shifts, while online hypothesis tests detected shifts with low false alarms. Memory classifiers and Persistence of Excitation (PoE) enhanced generalization and robust optimization. Additional advances included probabilistic invariance learning, confidence calibration, neuro-symbolic Multi-instance Partial Label Learning (PLL) with learnability guarantees, 68% accuracy in autism-enriched conversation quality prediction (surpassing human raters), multilingual LLM guardrails via synthetic data and GRPO, new benchmarks (NTSEBENCH, FlowVQA), and sanity checks for text/table/vision robustness. Validated across datasets and robots, these countered noise/texture-based attacks.
  • Third Approach: Compositional Inference for Runtime Adversaries
    Addressed concept-specific threats with block-world models defining robustness expectations and adversarial/compositional augmentation for atomic concept learning. Few-shot program-based learning outperformed neural parsers, while DFA state prediction excelled for regular languages, extending to vision-language navigation and NLP. Label smoothing, constrained regularization, and cross-modal attention in VLN boosted robustness. Evaluated in NLP, navigation, and robotics, these protected against partial compromises.

Thrust II: Adaptive Learning in Dynamic Environments

This thrust created adaptive frameworks to detect, repair, and lifelong-learn in changing conditions, enabling rapid response to novel attacks.

  • First Approach: Detecting/Fixing Inconsistencies with Inductive Biases
    Smaller models + label-smoothing improved out-of-domain calibration; semantic navigation estimated uncertainties for unseen areas. TAWT minimized cross-task distances (validated in NLP); incidental supervision and modular QA advanced NER/QA generalization. Memory-based shift detection in CPS (CARLA/LiDAR) provided interpretable alerts; MCNNs constrained imitation learning for manipulation/driving. SuperStAR recovered 14.21% accuracy on corrupted ImageNet-C/CIFAR-100-C via RL-optimized transforms. Data programming boosted medical labeling by 17–28% accuracy / 5–15% F1. AR-Pro, MINA, and others repaired/explained anomalies. Tested on robots, these supported fast adaptation.
  • Second Approach: Lifelong Concept Learning for Robustness/Adaptation Piaget-inspired reuse of concepts
    enhanced modularity in classification/RL. SHELS enabled OOD detection and incremental learning, extended to robotic segmentation. Compositional lifelong methods balanced stability-plasticity, accelerating RL tasks by 82.5% with neural modules. In collaboration with Julia Parish-Morris et al., “Sticky Mittens” (CoRL 2022 demo) applied infant-inspired RL with expert-guided option templates, improving exploration efficiency and safety. Option templates sped long-horizon RL; cross-modal VLN predicted waypoints; robust/replicable RL and IBCL/REGENT generalized across robotics/games. TransientTABLES advanced LLM temporal reasoning. Validated on robots, these enabled adaptation to evolving tasks.

Thrust III: Verification and Monitoring of Learning

This thrust delivered scalable tools to detect, verify, and monitor learned models in dynamic settings, ensuring trustworthiness.

VisionGuard outperformed on digital/physical adversarial attacks via transformations; Verisig gained scalability/precision; T4V trained verifiable properties. Video attack detection >85% accuracy; memory-based OOD offered interpretable CPS feedback. TRAQ reduced RAG hallucination sets by 16.2%; ISAR repaired controllers while preserving behaviors. PAC sets, MoCAN/PeCAN, Rank-Calibration, SocREval, BIRD, and others improved uncertainty, alignment, reasoning evaluation, and medical safety. Concurrent probes and conformal prediction monitored integrity and predicted STL violations. Validated on robots, these provided robust verification/monitoring.