Data Skeptic
Data Skeptic is a podcast hosted by Kyle Polich, featuring interviews with experts on data science, machine learning, AI, and statistics.
Seasons
- Recommender Systems — The're not just for eCommerce. Recommender systems have infiltrated our video watching, travel planning, social media feeds, and other daily activities. This se
- Graphs and Networks — Connections matter — in social media, biology, transportation, and beyond. This season maps out the study of graphs and networks, explaining core concepts like
- Animal Intelligence — How smart are animals — and how do we know? With a special co-host joining the show, this season explores cognition across the animal kingdom. Featuring fieldwo
- Machine Intelligence — What makes machines intelligent — and how close are we to achieving it? Through interviews and explainers, this season probes the current state of AI, exploring
- All About Surveys — Surveys might seem simple, but good data collection requires nuance, rigor, and constant vigilance. This season digs into the science behind asking questions: d
- Ad-tech — Online ads don’t just appear — they’re chosen through rapid-fire auctions, behavioral tracking, and complex optimization algorithms. These episodes reveal the m
- Physically Distributed — When systems and people are geographically separated, coordination becomes a technical and social challenge. This season investigates how distributed work and l
- k-Means Clustering — Clustering helps us find structure in unlabeled data — and this short season is a deep dive into one of the most popular algorithms: *k*-means. With clear expla
- Time Series — Understanding how data evolves over time is key to everything from forecasting demand to interpreting social trends. This season breaks down the structure and b
- Interpretability — As AI grows more powerful, understanding how models make decisions becomes critical. This season explores the tools and frameworks for interpreting machine lear
- Consensus — What does it mean to agree? This season explores how consensus is formed, whether among people or machines. Episodes dive into distributed systems protocols lik
- Natural Language Processing — Teaching computers to understand language is one of the most challenging problems in AI. This season walks through the evolution of natural language processing,
- Artificial Intelligence — AI is reshaping every industry, but what exactly is it — and what isn’t it? These episodes aim to define artificial intelligence in a rapidly evolving landscape
- Fake News — Misinformation isn't new, but the modern web has given it unprecedented reach and influence. This season investigates how false narratives spread, why they're s
All episodes
2026
- Recommender Systems Optimization Goals
- Recommender Systems Origin Story
- Social Choice for Fair Recommendations
- News Recommendations
- Give Users the Wheel
- AutoLike
- Student Spotlight: Aaron Payne, Data Analyst
- The Future is Agentic in Recommender Systems
- Book Ratings and Recommendations
- Disentanglement and Interpretability in Recommender Systems
- Collective Altruism in Recommender Systems
- Niche vs Mainstream
- Healthy Friction in Job Recommender Systems
- Fairness in PCA-Based Recommenders
2025
- Video Recommendations in Industry
- Eye Tracking in Recommender Systems
- Cracking the Cold Start Problem
- Designing Recommender Systems for Digital Humanities
- DataRec Library for Reproducible in Recommend Systems
- Shilling Attacks on Recommender Systems
- Music Playlist Recommendations
- Bypassing the Popularity Bias
- Sustainable Recommender Systems for Tourism
- Interpretable Real Estate Recommendations
- Why am I seeing This?
- Eco-aware GNN Recommenders
- Networks and Recommender Systems
- Network of Past Guests Collaborations
- The Network Diversion Problem
- Complex Dynamics in Networks
- Github Network Analysis
- Networks and Complexity
- Actantial Networks
- Graphs for Causal AI
- Power Networks
- Unveiling Graph Datasets
- Network Manipulation
- The Small World Hypothesis
- Thinking in Networks
- Fraud Networks
- Criminal Networks
- Graph Bugs
- Organizational Network Analysis
- Organizational Networks
- Networks of the Mind
- LLMs and Graphs Synergy
- A Network of Networks
- Auditing LLMs and Twitter
- Fraud Detection with Graphs
- Optimizing Supply Chains with GNN
- The Mystery Behind Large Graphs
2024
- Customizing a Graph Solution
- Graph Transformations
- Networks for AB Testing
- Lessons from eGamer Networks
- Github Collaboration Network
- Graphs and ML for Robotics
- Graphs for HPC and LLMs
- Graph Databases and AI
- Network Analysis in Practice
- Animal Intelligence Final Exam
- Process Mining with LLMs
- Open Animal Tracks
- Bird Distribution Modeling with Satbird
- Ant Encounters
- Computing Toolbox
- Biodiversity Monitoring
- Hacking the Colony
- Primate Poses
- Generating 3D Animals with YouDream
- Weird Communication
- Reducing the Impact of Ship Noise on Marine Mammals
- Analysis of Unstructured Data
- iNaturalist
- Learn to Code
- Animal Computer Interaction
- Ape Gestures
- Evaluating AI Abilities
- HMMs for Behavior
- Bioinspired Engineering
- Modelling Evolution
- Behavioral Genetics
- Signal in the Noise
- Pose Tracking
- Modeling Group Behavior
- Advances in Data Loggers
- What You Know About Intelligence is Wrong (fixed)
- Animal Decision Making
- Octopus Cognition
- Optimal Foraging
- Memory in Chess
- OpenWorm
- What the Antlion Knows
- AI Roundtable
2023
- Uncontrollable AI Risks
- I LLM and You Can Too
- Q&A with Kyle
- LLMs for Data Analysis
- AI Platforms
- Deploying LLMs
- A Survey Assessing Github Copilot
- Program Aided Language Models
- Which Programming Language is ChatGPT Best At
- GraphText
- arXiv Publication Patterns
- Do LLMs Make Ethical Choices
- Emergent Deception in LLMs
- Agents with Theory of Mind Play Hanabi
- LLMs for Evil
- The Defeat of the Winograd Schema Challenge
- LLMs in Social Science
- LLMs in Music Composition
- Cuttlefish Model Tuning
- Which Professions Are Threatened by LLMs
- Why Prompting is Hard
- Automated Peer Review
- Prompt Refusal
- A Long Way Till AGI
- Computable AGI
- AGI Can Be Safe
- AI Fails on Theory of Mind Tasks
- AI for Mathematics Education
- Evaluating Jokes with LLMs
- Why Machines Will Never Rule the World
- A Psychopathological Approach to Safety in AGI
- The NLP Community Metasurvey
- Skeptical Survey Interpretation
- The Gallup Poll
- Inclusive Study Group Formation at Scale
- The PhilPapers Survey
- Non-Response Bias
- Measuring Trust in Robots with Likert Scales
- CAREER Prediction
- The Panel Study of Income Dynamics
- Survey Design Working Session
- Bot Detection and Dyadic Surveys
- Reproducible ESP Testing
- A Survey of Data Science Methodologies
- Opinion Dynamics Models
- Casual Affective Triggers
- Conversational Surveys
- Do Results Generalize for Privacy and Security Surveys
- 4 out of 5 Data Scientists Agree
2022
- Crowdfunded Board Games
- Russian Election Interference Effectiveness
- Placement Laundering Fraud
- Data Clean Rooms
- Dark Patterns in Site Design
- Internet Advertising Bureau Media Lab
- Your Mouse Reveals Your Gender and Age
- Measuring Web Search Behavior
- StrategyQA and Big Bench
- Ad Blockers Effect on News Consumption
- Your Consent is Worth 75 Euros a Year
- Automated Email Generation for Targeted Attacks
- Tribal Marketing
- Debiasing GPT-3 Job Ads
- ML Ops in Production
- Ad Network Tomography
- First Party Tracking Cookies
- The Harms of Targeted Weight Loss Ads
- Podcast Advertising
- Fairness in e-Commerce Search
- Fraudulent Amazon Reviewers
- Ad Targeting in Amazon Smart Speakers
- Adwords with Unknown Budgets
- ML Ops Best Practices
- Affiliate Marketing Rabbithole
- Monetization of Youtube Conspiracy Theorists
- User Perceptions of Problematic Ads
- Political Digital Advertising Analysis
- Privacy Preference Signals
- Neural Architecture Search for CTR Prediction
- Algorithmic PPC Management
- Data Skeptic: Ad Tech
- The Reliability of Mobile Phone Data
- Haywire Algorithms
- School Reopening Analysis
- Modern Data Stacks
- Emoji as a Predictor
- Polarizing Trends in the Gig Economy
- Remote Learning in Applied Engineering
- Remote Productivity
- Does Remote Learning Work?
- Covid-19 Impact on Bicycle Usage
- Learning Digital Fabrication Remotely
- Remote Software Development
- Quantum K-Means
- K-Means in Practice
- Fair Hierarchical Clustering
- Matrix Factorization For k-Means
- Breathing K-Means
- Explainable K-Means
- Customer Clustering
- k-means Image Segmentation
- Tracking Elephant Clusters
- k-means clustering
- Snowflake Essentials
- Explainable Climate Science
- Energy Forecasting Pipelines
- Matrix Profiles in Stumpy
- The Great Australian Prediction Project
- Water Demand Forecasting
- Open Telemetry
2021
- Fashion Predictions
- Time Series Mini Episodes
- Forecasting Motor Vehicle Collision
- Deep Learning for Road Traffic Forecasting
- Bike Share Demand Forecasting
- Forecasting in Supply Chain
- Black Friday
- Aligning Time Series on Incomparable Spaces
- Comparing Time Series with HCTSA
- Change Point Detection Algorithms
- Time Series for Good
- Long Term Time Series Forecasting
- Fast and Frugal Time Series Forecasting
- Causal Inference in Educational Systems
- Boosted Embeddings for Time Series
- Change Point Detection in Continuous Integration Systems
- Applying k-Nearest Neighbors to Time Series
- Ultra Long Time Series
- MiniRocket
- ARiMA is not Sufficient
- Comp Engine
- Detecting Ransomware
- GANs in Finance
- Predicting Urban Land Use
- Opportunities for Skillful Weather Prediction
- Predicting Stock Prices
- N-Beats
- Translation Automation
- Time Series at the Beach
- Automatic Identification of Outlier Galaxy Images
- Do We Need Deep Learning in Time Series
- Detecting Drift
- Darts Library for Time Series
- Forecasting Principles and Practice
- Prequisites for Time Series
- Orders of Magnitude
- They're Coming for Our Jobs
- Pandemic Machine Learning Pitfalls
- Flesch Kincaid Readability Tests
- Fairness Aware Outlier Detection
- Life May be Rare
- Social Networks
- The QAnon Conspiracy
- Benchmarking Vision on Edge vs Cloud
- Goodhart's Law in Reinforcement Learning
- Video Anomaly Detection
- Fault Tolerant Distributed Gradient Descent
- Decentralized Information Gathering
- Leaderless Consensus
- Automatic Summarization
- Gerrymandering
- Consecutive Votes in Paxos
- Visual Illusions Deceiving Neural Networks
2020
- Earthquake Detection with Crowd-sourced Data
- Byzantine Fault Tolerant Consensus
- Alpha Fold
- Arrow's Impossibility Theorem
- Face Mask Sentiment Analysis
- Counting Briberies in Elections
- Sybil Attacks on Federated Learning
- Differential Privacy at the US Census
- Distributed Consensus
- ACID Compliance
- National Popular Vote Interstate Compact
- Defending the p-value
- Retraction Watch
- Crowdsourced Expertise
- The Spread of Misinformation Online
- Consensus Voting
- Voting Mechanisms
- False Consensus
- Fraud Detection in Real Time
- Listener Survey Review
- Human Computer Interaction and Online Privacy
- Authorship Attribution of Lennon McCartney Songs
- GANs Can Be Interpretable
- Sentiment Preserving Fake Reviews
- Interpretability Practitioners
- Facial Recognition Auditing
- Robust Fit to Nature
- Black Boxes Are Not Required
- Robustness to Unforeseen Adversarial Attacks
- Estimating the Size of Language Acquisition
- Interpretable AI in Healthcare
- Understanding Neural Networks
- Self-Explaining AI
- Plastic Bag Bans
- Self Driving Cars and Pedestrians
- Computer Vision is Not Perfect
- Uncertainty Representations
- AlphaGo, COVID-19 Contact Tracing and New Data Set
- Interpretability Tooling
- Shapley Values
- Anchors as Explanations
- Adversarial Explanations
- ObjectNet
- Visualization and Interpretability
- Interpretable One Shot Learning
- Fooling Computer Vision
- Algorithmic Fairness
- Interpretability
2019
- NLP in 2019
- The Limits of NLP
- Jumpstart Your ML Project
- Serverless NLP Model Training
- Team Data Science Process
- Ancient Text Restoration
- ML Ops
- Annotator Bias
- NLP for Developers
- Indigenous American Language Research
- Talking to GPT-2
- Reproducing Deep Learning Models
- What BERT is Not
- SpanBERT
- BERT is Shallow
- BERT is Magic
- Applied Data Science in Industry
- Building the howto100m Video Corpus
- BERT
- Onnx
- Catastrophic Forgetting
- Transfer Learning
- Facebook Bargaining Bots Invented a Language
- Under Resourced Languages
- Named Entity Recognition
- The Death of a Language
- Neural Turing Machines
- Data Infrastructure in the Cloud
- The Transformer
- Mapping Dialects with Twitter Data
- Sentiment Analysis
- Attention Primer
- Cross-lingual Short-text Matching
- ELMo
- BLEU
- Simultaneous Translation at Baidu
- Human vs Machine Transcription
- seq2seq
- Text Mining in R
- Recurrent Relational Networks
- Text World and Word Embedding Lower Bounds
- word2vec
- Authorship Attribution
- Very Large Corpora and Zipf's Law
- Semantic search at Github
2018
- Data Science Hiring Processes
- Holiday Reading - Epicac
- Drug Discovery with Machine Learning
- Sign Language Recognition
- Data Ethics
- Escaping the Rabbit Hole
- [MINI] Theorem Provers
- Automated Fact Checking
- [MINI] Single Source of Truth
- Detecting Fast Radio Bursts with Deep Learning
- Being Bayesian
- Modeling Fake News
- The Louvain Method for Community Detection
- Cultural Cognition of Scientific Consensus
- False Discovery Rates
- Deep Fakes
- Fake News Midterm
- Quality Score
- The Knowledge Illusion
- Click Through Rates
- Algorithmic Detection of Fake News
- Ant Intelligence
- Human Detection of Fake News
- Spam Filtering with Naive Bayes
- The Spread of Fake News
- Fake News
- Dev Ops for Data Science
- First Order Logic
- Blind Spots in Reinforcement Learning
- Defending Against Adversarial Attacks
- Transfer Learning
- Medical Imaging Training Techniques
- Kalman Filters
- AI in Industry
- AI in Games
- Game Theory
- The Experimental Design of Paranormal Claims
- Winograd Schema Challenge
- The Imitation Game
- Eugene Goostman
- The Theory of Formal Languages
- The Loebner Prize
- Chatbots
- The Master Algorithm
- The No Free Lunch Theorems
- ML at Sloan Kettering Cancer Center
- Optimal Decision Making with POMDPs
- AI Decision-Making
- [MINI] Reinforcement Learning
- Evolutionary Computation
- [MINI] Markov Decision Processes
- Neuroscience Frontiers
- Neuroimaging and Big Data
- The Agent Model of Artificial Intelligence
2017
- Artificial Intelligence, a Podcast Approach
- Holiday reading 2017
- Complexity and Cryptography
- Mercedes Benz Machine Learning Research
- [MINI] Parallel Algorithms
- Quantum Computing
- Azure Databricks
- [MINI] Exponential Time Algorithms
- P vs NP
- [MINI] Sudoku \in NP
- The Computational Complexity of Machine Learning
- [MINI] Turing Machines
- The Complexity of Learning Neural Networks
- [MINI] Big Oh Analysis
- Data science tools and other announcements from Ignite
- Generative AI for Content Creation
- [MINI] One Shot Learning
- Recommender Systems Live from FARCON 2017
- [MINI] Long Short Term Memory
- Zillow Zestimate
- Cardiologist Level Arrhythmia Detection with CNNs
- [MINI] Recurrent Neural Networks
- Project Common Voice
- [MINI] Bayesian Belief Networks
- pix2code
- [MINI] Conditional Independence
- Estimating Sheep Pain with Facial Recognition
- CosmosDB
- [MINI] The Vanishing Gradient
- Doctor AI
- [MINI] Activation Functions
- MS Build 2017
- [MINI] Max-pooling
- Unsupervised Depth Perception
- [MINI] Convolutional Neural Networks
- Multi-Agent Diverse Generative Adversarial Networks
- [MINI] Generative Adversarial Networks
- Opinion Polls for Presidential Elections
- OpenHouse
- [MINI] GPU CPU
- [MINI] Backpropagation
- Data Science at Patreon
- [MINI] Feed Forward Neural Networks
- Reinventing Sponsored Search Auctions
- [MINI] The Perceptron
- The Data Refuge Project
- [MINI] Automated Feature Engineering
- Big Data Tools and Trends
- [MINI] Primer on Deep Learning
- Data Provenance and Reproducibility with Pachyderm
- [MINI] Logistic Regression on Audio Data
- Studying Competition and Gender Through Chess
- [MINI] Dropout
- The Police Data and the Data Driven Justice Initiatives
2016
- The Library Problem
- 2016 Holiday Special
- [MINI] Entropy
- MS Connect Conference
- Causal Impact
- [MINI] The Bootstrap
- [MINI] Gini Coefficients
- Unstructured Data for Finance
- [MINI] AdaBoost
- Stealing Models from the Cloud
- [MINI] Calculating Feature Importance
- NYC Bike Share Rebalancing
- [MINI] Random Forest
- Election Predictions
- [MINI] F1 Score
- Urban Congestion
- [MINI] Heteroskedasticity
- Music21
- [MINI] Paxos
- Trusting Machine Learning Models with LIME
- [MINI] ANOVA
- Machine Learning on Images with Noisy Human-centric Labels
- [MINI] Survival Analysis
- Predictive Models on Random Data
- [MINI] Receiver Operating Characteristic (ROC) Curve
- Multiple Comparisons and Conversion Optimization
- [MINI] Leakage
- Predictive Policing
- [MINI] The CAP Theorem
- Detecting Terrorists with Facial Recognition?
- [MINI] Goodhart's Law
- Data Science at eHarmony
- [MINI] Stationarity and Differencing
- Feather
- [MINI] Bargaining
- deepjazz
- [MINI] Auto-correlative functions and correlograms
- Early Identification of Violent Criminal Gang Members
- [MINI] Fractional Factorial Design
- Machine Learning Done Wrong
- Potholes
- [MINI] The Elbow Method
- Too Good to be True
- [MINI] R-squared
- Models of Mental Simulation
- [MINI] Multiple Regression
- Scientific Studies of People's Relationship to Music
- [MINI] k-d trees
- Auditing Algorithms
- [MINI] The Bonferroni Correction
- Detecting Pseudo-profound BS
- [MINI] Gradient Descent
- Let's Kill the Word Cloud
2015
- 2015 Holiday Special
- Wikipedia Revision Scoring as a Service
- [MINI] Term Frequency - Inverse Document Frequency
- The Hunt for Vulcan
- [MINI] The Accuracy Paradox
- Neuroscience from a Data Scientist's Perspective
- [MINI] Bias Variance Tradeoff
- Big Data Doesn't Exist
- [MINI] Covariance and Correlation
- Bayesian A/B Testing
- Let's Talk About Natural Language Processing