Posts

Showing posts from September, 2021

Video Encoders

 Is Space-Time Attention All You Need for Video Understanding? A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

Explainable AI

 From Probability to Consilience:How Explanatory Values Implement Bayesian Reasoning On quantitative aspects of model interpretability Cognitive Perspectives on Context-based Decisions and Explanations From Human Explanation to Model Interpretability:A Framework Based on Weight of Evidence Manipulating and Measuring Model Interpretability Interpretable Machine Learning: FundamentalPrinciples and 10 Grand Challenges

Document Machine Translation

 Measuring and Increasing Context Usage inContext-Aware Machine Translation

Numerical Reasoning

 Injecting Numerical Reasoning Skills into Language Models

Algebra Word Problems

 Learning to Automatically Solve Algebra Word Problems

Vision-Language Navigation

 The Road to Know-Where: An Object-and-Room Informed Sequential BERTfor Indoor Vision-Language Navigation Vision-Language Navigation with Random Environmental Mixup

Speech Translation

 LEVERAGING WEAKLY SUPERVISED DATA TO IMPROVE END-TO-ENDSPEECH-TO-TEXT TRANSLATION Beyond Sentence-Level End-to-End Speech Translation: Context Helps

Speech Pre-training

  W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training Injecting Text in Self-Supervised Speech Pretraining Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification EAT: Enhanced ASR-TTS for Self-supervised Speech Recognition

Energy-based Models

 RESIDUAL ENERGY-BASED MODELS FOR TEXT GENERATION

Content Planning in Text Generation

 DYPLOC: Dynamic Planning of Content Using Mixed Language Modelsfor Text Generation

Non/Semi Auto-regressive Text Generation

 Cascaded Text Generation with Markov Transformers Limitations of Autoregressive Models and Their Alternatives

Deep Structured Prediction

 Torch-Struct:Deep Structured Prediction Library Structured Prediction Cascades

Multimodal Pretrained Models

 AudioCLIP: Extending CLIP to Image, Text and Audio

Neural Grammar Learning

 Visually Grounded Compound PCFGs Sequence-to-Sequence Learning with latent Neural Grammars

Scene Graphs

 Scene Graph Parsing as Dependency Parsing

Continual Learning

 Refining Sample Embeddings with Relation Prototypes toEnhance Continual Relation Extraction Tong's literature on github

Causality

 Introduction to Judea Pearl’s Do-Calculus

Scalable Transformers

 SWITCH TRANSFORMERS: SCALING TO TRILLIONPARAMETER MODELS WITH SIMPLE AND EFFICIENTSPARSITY Deep Speed GPT-J

Small Language Models Are Also Few-Shot Learners

 It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners Big Self-Supervised Models areStrong Semi-Supervised Learners

Event Extraction as Generation with Language Models

TEXT2EVENT: Controllable Sequence-to-Structure Generation for End-to-end Event Extraction  DEGREE: A Data-Efficient Generative Event Extraction Model

Transformers and Circuit Complexity (tbf)

 On the Power of Saturated Transformers:A View from Circuit Complexity

Multimodal Conversational Agents (TBF)

Maria: A Visual Experience Powered Conversational Agent Neural Abstructions: Abstractions that SupportConstruction for Grounded Language Learning

Neural Transducers (TBF)

 TRANSFORMER TRANSDUCER: A STREAMABLE SPEECH RECOGNITION MODEL WITH TRANSFORMER ENCODERS AND RNN-T LOSS An Online Sequence-to-Sequence Model Using PartialConditioning Sequence Transduction with Recurrent Neural Networks

Data Augmentation for Low-resource Semantic Parsing

 AutoQA: From Databases To QA Semantic ParsersWith Only Synthetic Training Data

Understanding Self-supervised Learning (TBF)

 Understanding Self-Supervised Learning Dynamics without Contrastive Pairs

Random Feature Attention (TBF)

 Random Feature Attention   [paper] [slide]

Encoder-Decoder Significance in NMT (TBF)

 Analyzing the Source and Target Contributions to Predictions in Neural Machine Translation DEEP ENCODER, SHALLOW DECODER: REEVALUATING NON-AUTOREGRESSIVE  MACHINE TRANSLATION Multilingual Neural Machine Translation withDeep Encoder and Multiple Shallow Decoders

Information Bottleneck (TBF)

 VARIATIONAL INFORMATION BOTTLENECK FOR EFFECTIVE LOW-RESOURCE FINE-TUNING

LLM fine-tuning (TBF)

 REVISITING A FEW-SAMPLE BERT FINE-TUNING Recent Advances in Language Model Fine-tuning ON THE STABILITY OF FINE-TUNING BERT: MISCONCEPTIONS, EXPLANATIONS, AND STRONG BASELINES

Meta Learning

 Online Structured Meta-learning Reptile: A Scalable Meta-Learning Algorithm META-DATASET: A DATASET OF DATASETS FORLEARNING TO LEARN FROM FEW EXAMPLES

Prompting in NLU and NLG (TBF)

 Pre-train, Prompt, and Predict: A Systematic Survey ofPrompting Methods in Natural Language Processing

Semantic Parsing for KBs

 TURING: an Accurate and Interpretable Multi-Hypothesis Cross-DomainNatural Language Database Interface

Bi-level Optimisation

 Rethinking Bi-Level Optimization in Neural Architecture Search:A Gibbs Sampling Perspective Investigating Bi-Level Optimization for Learning and Vision from a Unified Persp Bilevel Programming for Hyperparameter Optimization and Meta-Learning

Internet-Augmented Dialogue Generation

 Internet-Augmented Dialogue Generation

Sub-Networks in NMT (TBF)

 Importance-based Neuron Allocation for Multilingual Neural Machine Translation  Parameter-Efficient Transfer Learning with Diff Pruning

Multi-Source Domain Adaptation

 MOST: Multi-Source Domain Adaptation via Optimal Transport for Student-Teacher Learning On Deep Domain Adaptation: Some TheoreticalUnderstandings

Perceiver and Perceiver IO (Deepmnind)

Deepmind's Blog Post. These papers suggest efficient Transformer architecures for long inputs. This provides a powerful backbone for handling data from different modalities as well in a seamless manner. The model only looks at 'arrays of bytes' without regard to the nature of the data (eg image or text). Perceiver IO: A General Architecture for Structured Inputs & Outputs

Zero-Shot Text-to-Image Generation (TBF)

 Zero-Shot Text-to-Image Generation

Adapters in NLU and NLG

Parameter-efficient Multi-task Fine-tuning for Transformersvia Shared Hypernetworks

Text-based Games (TBF)

 Modeling Worlds in Text

Compressing Encoder (TBF)

 On Sparsifying Encoder Outputs in Sequence-to-Sequence Models Learned Token Pruning for Transformers

GRADIENT REGULARIZATION in DNNs

 GRADIENT REGULARIZATION IMPROVES THE ACCURACY OF DISCIMINATIVE MODELS A UNIFYING VIEW ON IMPLICIT BIAS IN TRAININGLINEAR NEURAL NETWORKS

Emergent linguistic structure in NNs & Understanding MLMs (TBF)

  Emergent linguistic structure in artificial neural networks trained by self-supervision On the Inductive Bias of Masked Language Modeling: From Statistical to Syntactic Dependencies A Consistency Theorem for BERT

Time-Stamped Language Model (TBF)

  Time-Stamped Language Model: Teaching Language Models to understand the Flow of Events

Compositional Question Answering

  End-to-End Multihop Retrieval for Compositional Question Answering over Long Documents Unsupervised Question Decomposition for Question Answering CANDLE: Decomposing Conditional and Conjunctive Queries for Task-Oriented Dialogue Systems Break, Perturb, Build: Automatic Perturbation of Reasoning Paths through Question Decomposition Natural Logic  (1970 paper) Stanford's page on natural logic

PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World (TBF)

  PIGLeT: Language Grounding Through Neuro-Symbolic Interaction in a 3D World

Knowledge in Large Pretrained Language Models and Neural DBs (TBF)

  MERLOT: Multimodal Neural Script Knowledge Models Neural Databases Database Reasoning Over Text X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models How Context Affects Language Models’ Factual Predictions How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering

Deep Architectures for Reasoning (TBF)

  Understanding Deep Architectures with Reasoning Layer  (Le Song) Transformers as Soft Reasoners over Language NLProlog: Reasoning with Weak Unification for question Answering in Natural Language From Deep Learning to Deep Reasoning DEEP LEARNING NEEDS A PREFRONTAL CORTEX  (ICLR2020, Bengio)  Machine Learning to Machine Reasoning Self-Attentive Associative Memory

Language Models and Zero-shot Learning (TBF)

  LANGUAGE MODELS ARE ZERO-SHOT LEARNERS Language Models are Few-Shot Learners

Hierarchical Knowledge Distillation

  META-LEARNING FOR KNOWLEDGE DISTILLATION  (Julian McAluey) Xuanli's report

Self-supervised Learning for Reasoning and Perception (TBF)

  The workshop page (ICML2011)

Constrained Decoding and Controlled Generation (TBF; idea?)

 TBF NEUROLOGIC DECODING:(Un)supervised Neural Text Generation with Predicate Logic Constraints Zero-Shot Controlled Generation with Encoder-Decoder Transformers Natural Language Generation & Evaluation (Asli's talk) THE CURIOUS CASE OF NEURAL TEXT DeGENERATION (nucleus sampling, ICLR 2020)

Natural Theorem Proving (TBF; idea?)

TBF  NATURALPROOFS: Mathematical Theorem Provingin Natural Language

Commonsense Reasoning and KG (TBF)

 TBF Conversational Neuro-Symbolic Commonsense Reasoning CommonsenseQA 2.0: Exposing the Limits of through Gamification TIMEDIAL: Temporal Commonsense Reasoning in Dialog

Semantics in Neural Machine Translation (TBF; idea?)

  TBF Semantic Neural Machine Translation using AMR

Compositional Generalisation (TBF)

TBF semantic parsing Improving Compositional Generalization in Semantic Parsing Compositional Generalization for Neural Semantic Parsing via Span-level Supervised Attention Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional Generalization M EASURING C OMPOSITIONALITY IN R EPRESENTATION L EARNING machine translation On Compositional Generalization of Neural Machine Translation ML COMPOSITIONAL GENERALIZATION WITH TREE STACK MEMORY UNITS Generalization without Systematicity:On the Compositional Skills of Sequence-to-Sequence Recurrent Networks

Scaling Laws for Transformers (TBF)

TBF  SCALE EFFICIENTLY: INSIGHTS FROM PRE-TRAINING AND FINE-TUNING TRANSFORMERS Scaling Laws for Neural Machine Translation Scaling laws for neural language  models Scaling laws for autoregressive  generative modeling

Focused Presentations on the Literature

 I have started to put together presentations on focused topics. My hope is that they can provide structure to the literature of various research directions that we pursue in our group.  To make them accessible easily, I provide the links below for future reference:  1) Deep Transfer Learning (adapters, subnetwork, etc) LINK 2) Commonsense Reasoning (Yejin) LINK   3) Semantic Parsing (Jonathan) LINK 4) Open-Domain Question Answering (Hannaneh, Danqi) LINK 5) Neural Speech Recognition (pretraining etc) LINK 6) Neural Speech Translation (simultaneous, context, etc)  LINK 7) Neural Machine Translation (context, simultaneous, low resource, etc) LINK These presentations can be used for the reading group, and they can be gradually updated as new papers are published.  Relatedly, sanity arxiv can be used to get the most recent work related to your research from arxiv, rather than being bombarded by the flood of new papers published every day. The papers can be fi...

Humans Learn From Task Descriptions and So Should Our Models

Image
An excellent talk on PE T: https://www.youtube.com/watch?v=_YOaRLQnBjc&t=246s TDLR; They provide a task description (patterns + class label verbalizer; abbreviated by PET) plus a few examples of the target task of interest to a pre-trained LLM, and train the LLM based on the few examples. Their main difference to GPT3 prompting is that they allow the model to be trained based on the few examples in the prompt. They show that it leads to big improvements compared to GPT3, even if the size of their LLM is much smaller compared to GPT3. The intuition is that, as humans, we update our brains and 'learn' after receiving examples. Can ‘one’ model be trained for a large number of patterns? Yes, maybe in a sequential manner. That is, the problem is now reduced to ‘continual learning', or even multitask learning? Maybe subnetwork pruning can help here, to identify the relevant subnetwork and only update those parameters.   Can the unlabeled data be used to filter our good patte...