DS Stream logo
AI Transformation
AI TRANSFORMATION
Services
Services
SERVICES
Solutions
Solutions
SOLUTIONS
Insights
INSIGHTS
Blog
Case Studies
Webinars
About Us
ABOUT US
CONTACT
Blog
Michal Milosz
Latest blog posts by Michal Milosz
Latest blog posts
View all
BI and Big Data Consulting
Custom Software Development
DevOps Managed Services
Data Pipelines
AI Agents
AI in Retail
AI Engineering
AI Transformation
AI Sales
AI Security
Generative AI
Databricks
CSR
MLOps
Google Cloud Platform
Data Migration
Data Analysis
Data Engineering
DevOps
Quantum Computing
Augmented Reality
Internet of Things
Blockchain Technology
Cyber Security
Cloud Computing
Data Science
Artificial Intelligence
Machine Learning
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
No blog posts found.
Databricks
9
min read
Managing Large Data Sets in Databricks
Optimize large datasets in Databricks with Partitioning, Z-Ordering, Auto Optimize, Delta Lake Vacuum, Caching, and Cost Monitoring for better performance.
Read more
Data Engineering
12
min read
Configuring the Celery Kubernetes
Configure the Celery Kubernetes Executor in Airflow 2.0 to combine Celery's scalability with Kubernetes' resource optimization for efficient workflows.
Read more
Data Engineering
7
min read
The Future of Data Engineering – Trends to Watch in 2025
Discover the top data engineering trends for 2025, including AI-driven automation, Lakehouse architecture, serverless and edge computing, and sustainable data practices. Embracing these innovations will help organisations optimise data operations and stay ahead in a rapidly changing landscape.
Read more
Data Engineering
6
min read
Data Warehouse vs. Data Lake vs Lakehouse: A Comprehensive Comparison of Data Management Approaches
The article compares three key data architectures: data warehouses for structured analytics, data lakes for flexible raw data storage, and lakehouse as a unified model. It highlights their strengths, limitations, and when each is the best fit for modern organizations.
Read more
DevOps
9
min read
How to spin up to 1000 parallel Airflow 2.0 tasks in 5 minutes from nothing with the CeleryExecutor and Kubernetes
Read more
Data Science
18
min read
Stream Processing vs. Batch Processing - A Practical Guide to Data Processing
Explore stream vs batch processing, their benefits, use cases, and technologies. Learn how to choose the right data strategy for your business needs.
Read more
Data Engineering
13
min read
Microservices in Data Engineering: How to Break a Monolith into Smaller Parts
Discover how microservices revolutionize data engineering by replacing monolithic systems, boosting scalability, flexibility, and deployment efficiency.
Read more
Data Science
11
min read
From Excel to Data Lake: The Evolution of Data Storage in a Modern Organization
Explore the shift from Excel to data lakes, modern storage solutions, and best practices for scalable, efficient, and innovative data management strategies.
Read more
Data Science
12
min read
Lakehouse Federation in Databricks: A Practical Guide
Lakehouse Federation in Databricks: query external data sources without moving data. See how it works and when to use it in your data platform.
Read more
Previous
1