City Course Page Acad ID: ACAD0138
Data Analysis with Python and Dask Training in Washington, D.C., United States

Python and Dask training is designed to help participants perform scalable data analysis on large datasets using Python and Dask.

Overview

This Data Analysis with Python and Dask training is designed to help participants perform scalable data analysis on large datasets using Python and Dask. The course covers core data analysis concepts with Pandas, transitioning to Dask for parallel and distributed computing, handling large datasets that exceed memory, optimizing performance, and building efficient data pipelines. Participants will gain hands-on experience to analyze, process, and scale data workflows for modern analytics and data engineering use cases.

Learning Outcomes

• Understand data analysis fundamentals and build efficient data processing workflows using Python.
• Clean, transform, and analyze structured datasets using Pandas and NumPy.
• Perform exploratory data analysis and generate meaningful business insights from data.
• Process large-scale datasets efficiently using Dask for parallel and distributed computing.
• Create data visualizations and reports using Python visualization libraries.
• Build scalable data analysis pipelines for real-world analytics use cases.

Duration & Delivery Mode

14 hours

We serve:
Target Audience

 • Data analysts and business intelligence professionals
 • Data engineers and analytics engineers
 • Python developers working with large datasets
 • Data scientists handling big data workloads
 • Professionals transitioning to scalable analytics

Pre-requisites

 • Basic knowledge of Python programming
 • Familiarity with Pandas is helpful
 • Understanding of basic data analysis concepts

Skillset Achieved

 • Performing data analysis using Python and Pandas
 • Scaling data processing with Dask
 • Working with large datasets beyond memory limits
 • Building parallel and distributed data workflows
 • Optimizing data processing performance
 • Creating scalable data pipelines
 • Applying best practices for big data analytics

Course Outcome

By the end of this training, participants will be able to perform scalable data analysis using Python and Dask with confidence. Learners will understand how to transition from single-machine data processing to parallel and distributed workflows, enabling them to analyze large datasets efficiently and build production-ready analytics pipelines.

Course Outline

Foundations of Data Analysis with Python
 • Data analysis workflow overview
 • Reviewing Pandas for data manipulation
 • Loading and exploring datasets
 • Data cleaning and preprocessing basics

Introduction to Dask
 • What is Dask and when to use it
 • Dask vs Pandas and NumPy
 • Setting up Dask environment
 • Understanding lazy evaluation

Dask DataFrames & Arrays
 • Working with Dask DataFrames
 • Converting Pandas to Dask
 • Chunking and partitioning concepts
 • Performing basic transformations

Exploratory Data Analysis at Scale
 • Descriptive statistics with Dask
 • Filtering and aggregations
 • GroupBy operations
 • Handling missing values at scale

Parallel Computing & Task Scheduling
 • Dask task graphs
 • Parallel execution concepts
 • Local vs distributed schedulers
 • Monitoring task execution

Scaling Machine Learning Data Preparation
 • Feature engineering with Dask
 • Data normalization and encoding
 • Preparing large datasets for ML
 • Integrating with popular ML libraries

Performance Optimization Techniques
 • Partition sizing strategies
 • Memory management
 • Persist and cache operations
 • Diagnosing performance bottlenecks

Building Scalable Data Pipelines
 • Designing end-to-end data workflows
 • ETL pipeline concepts with Dask
 • Scheduling and automation basics
 • Best practices for production data pipelines

Deployment & Production Considerations
 • Running Dask on cloud and clusters
 • Security and access considerations
 • Logging and monitoring
 • Cost and resource optimization

Assessment Topics

• Python Data Handling & Preprocessing Assessment
• Data Cleaning & Transformation Assessment
• Exploratory Data Analysis & Visualization Assessment
• Distributed Data Processing with Dask Assessment
• End-to-End Analytics Mini Project Assessment

Evaluation

Participants will be evaluated through hands-on labs, real-world data analysis exercises, instructor-led workflow reviews, and a final practical assessment focused on building scalable Dask-powered data analysis solutions.

Course Materials

Participants will receive course materials, slides, reference materials, exercises and access to resources for further learning.

Certification

Upon successful completion of the training, participants will receive an AcadNXT Certificate of Completion for Data Analysis with Python and Dask. This digital, verifiable certification validates scalable data analytics and distributed computing skills and can be shared on LinkedIn and included in professional portfolios.

SELECT AN UPCOMING CLASS
Sat 15th Aug 2026 – Sun 16th Aug 2026
⏱ 2 days 📍 Classroom
AcadNXT Classroom - Washington, D.C Washington, D.C. United States
Sat 12th Sep 2026 – Sun 13th Sep 2026
⏱ 2 days 📍 Classroom
AcadNXT Classroom - Washington, D.C Washington, D.C. United States
No upcoming classes are currently available for this delivery mode.

Other cities in United States

Explore the same course in other cities across United States.

Back to United States course page

Enroll Now

WHO WILL BE FUNDING THE COURSE?

By submitting your details you agree to be contacted in order to respond to your enquiry.