Wetterdienst: Fast, Unified Access to Open Weather Data with Polars

In this presentation, Benjamin Gutzmann, a Data Engineer at Otto Group data.works, introduces Wetterdienst, a Python library designed to simplify the complex process of accessing open weather data. Because weather and environmental APIs vary wildly in format and structure, data engineers often spend more time on "plumbing" than on actual analysis. Benjamin explains how Wetterdienst solves this by providing a unified, Polars-first interface that standardizes request patterns across multiple global services, including the DWD, NOAA/NWS, and ECCC.

Viewers will learn how the library normalizes inconsistent data into tidy, long-format DataFrames using SI units and UTC timestamps, while implementing robust caching and retry mechanisms to ensure pipeline reliability. Benjamin walks through the provider architecture and demonstrates practical workflows for station discovery, timeseries retrieval, and exporting data to databases. Whether you are building ETL pipelines or training ML models, this talk provides a blueprint for integrating weather data via Python, CLI, or REST API to accelerate your analytics and operations.

This description was generated by Open-Source AI using the transcript of the session and the original submission contents.

This session took place in track Data Handling & Data Engineering and was classified suitable for novice domain / intermediate python by the speaker.

Submission

The proposal as submitted by the speaker before the conference.

Problem

Accessing weather data means wrestling with inconsistent APIs, formats, and units—slowing down data engineering and making pipelines hard to reproduce.

Solution

Wetterdienst is a Python library providing a unified, Polars-first interface to multiple open weather services (DWD, ECCC, EA, NOAA/NWS, Geosphere Austria, IMGW, Eaufrance, WSV, and more). It standardizes request patterns, returns tidy long-format data in SI units, and handles caching, timezones, and retries—so teams can focus on analysis instead of plumbing.

Core concepts:

  • Polars-first — All data operations use Polars (v1.15+); pandas supported for some I/O
  • Declarative request pattern — Provider → stations → values; tidy/long output by default
  • Sensible defaults — UTC timestamps, SI units, humanized parameter names
  • Reliability — Disk-based caching via diskcache, stamina-based retries, timezone handling
  • Provider architecture — Consistent interfaces across DWD, ECCC, EA, NOAA/NWS, Geosphere, IMGW, Eaufrance, WSV, and more
  • Multiple interfaces — Python API, CLI, and REST

Outline

  • Introduction
  • Journey — How Wetterdienst came to life
  • Wetterdienst — Architecture, concepts, and request patterns
  • Value — What wetterdienst offers you, me and everyone else
  • Demo — Live: station discovery, timeseries retrieval, station metadata, climate stripes and more via app

Target Audience

Data engineers, scientists, and platform teams who need reliable weather data for analytics, ML, and operations.

Prerequisites

Basic Python and DataFrame experience (Polars or pandas); familiarity with ETL/ML pipelines helpful.

Key Takeaways

  • A unified, Polars-first workflow to access and normalize open weather data
  • Practical patterns for station discovery, timeseries retrieval, unit conversion, and caching
  • How to integrate Wetterdienst via Python, CLI, and REST, and export to common formats and databases

📦 Repo https://github.com/earthobservations/wetterdienst 📖 Docs https://wetterdienst.readthedocs.io/ 🌐 App https://wetterdienst.eobs.org/ 💡 Examples https://github.com/earthobservations/wetterdienst/tree/main/examples

Transcript (auto)

Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.

Speaker 1 [00:06]

So, hello, welcome everybody. Actually, I just added the title itself this morning, because I forgot about it. And also there's another slogan we use usually, which is Open Earth Data for Humans. So Wetterdeans these days also covers not only weather data, but also some hydrological data. And maybe the slogan resonates with some of you from another library. No? We actually took it from requests, which uses HTTP for humans. So the idea is just to make it as simple as possible to retrieve weather and also Earth data and a few lines of code. And yeah, Wetterdeans also stands for Weather Service. And also a little warning, I will use the word DVD, which stands for German Weather Service. And I think the English pronunciation of DVD is just, it's not good. So I will skip that. Here's a little introduction. So first I'm going to introduce myself again a little bit. I'm talking about the journey of Wetterdeans or to Wetterdeans. Then about Wetterdeans and its features itself. Then the value that it gives to you, me and to everyone, hopefully. Then we have a little demo and afterwards the questions. So, I'm 32 years old. I'm a happy resident of Hamburg. I have several hobbies. So I usually play football, but I also like to go to Hamburg's biggest football stadium. I play also volleyball and also by day I like to walk Otto on the right side. And by night I like to go to cinema and concerts, for example, in the famous Mojo Club on the Ripperbahn. But obviously I also do coding by night when there's some time left. Yeah. I studied hydrology at Theodrasten. So I'm from the environmental sector. So from 2013 until 2020. When people ask me what is hydrology, I first tell them water. And the second thing I'm telling them is it's basically the entire water cycle. So from the clouds in the sky up to the bottom to the groundwater. Yeah. And why Dresden? Dresden has a long history with severe floodings. You see in the middle a picture of Dresden in 2002. The entire city centre was flooded. And on the right you see from the same year a picture of the main station. Imagine what conditions you had during this period. So I think that's a big relation between the city and hydrology. And in this case of flooding. Yeah. Since 2022 I'm at Otto Group DataWorks, which is also part of 1.0. We are also here with, I just put up the picture on the right. We are here with a big group of people. There are many young people also there. When I joined them it felt a bit like just being at university again. So not a real change. Yeah. And at the Otto Group I'm working usually on rankings for shops. But also these days we are working on generative AI based bots. which I'm feeling really lucky about. Because we can use all the nice technology. And we also on Google Cloud. So we have access to all the features of the Google Cloud. And it's a real pleasure to work there. And you get to learn lots of stuff. So then about the journey to Wetterdienst. I'm starting at how you can work as a hydrologist. There are many different possibilities. For example on the top left some people like to go in the field. Do some measurements on rivers like velocity and chemicals. And other people also go into simulation. For example on the bottom left the groundwater simulation is an important part. But others may go to some companies that manage reservoirs. And do some predictions and forecasting for the capacity. And on the bottom right some people tend to stick to university or to some research. And actually this is the Institute of Meteorology. And this is also where I started with my bachelor thesis. Which was about homogeneity. So homogeneity in terms of precipitation time series. So I was tasked to research the homogeneity of Eastern Saxony, Eastern Saxonian precipitation time series. And homogeneity in this case means, so usually you have your precipitation station. And you do the measurements all the time. But then maybe over decades, over hundreds of years, climate may change. And this could cause a disruption in the time series. So it may jump or it may, yeah, there may be a trend upwards. But there could also be external cases or like external changes in the surroundings. Like for example, a house is being built. The station could be moved 100 meters further. And then suddenly you have other conditions there. And precipitation may be also different because of those different conditions. And you don't, you want to make sure that when you say precipitation changes in this area. It's not because of the surroundings changes, but because of actual climate changes. So, and this was, this had to be done for, like I said, Eastern Saxony with roughly, I think 95 stations it was. And you could imagine this is not being done by Excel, but instead, and this is a, I think a typical entry to scripting and coding in the university. It's not Python, but I think it's R rather, because it's really simple. You have your script, you can run line by line. You can, I think it's easier to understand for the beginners. So this is how I started. So I was not into Python back then. But my interest was sparked and I continued working in the university. And then I thought, hmm, I have the opportunity as a student to join some conference for cheap money. So I went to Brussels in 2017, which was really nice. And there was also a talk about AirDVD or RDWD by a colleague who I think worked in Potsdam Climate Research Institute back then. And it felt really nice, like you have really only a few lines of code to achieve some data, which I think basically back when I worked in the institute, they did still some transfer of, like from, I think they, I don't know if they sent them hard drives from the DVD to the institute. And this, this library would just make it so much easier to receive the data. And in the same time, I thought, um, R is not enough for me. I think, um, you can do some stuff with it, but Python has so much more capabilities. So I want to learn Python and I have already a library that I like, um, and I need a project which I can use to learn Python. So I'm just gonna port it over to Python. Um, yeah, and that's, that's how it actually went. Um, along the way, I want to mention, uh, two important things. One is my mentor, I call him. I don't know if he, I don't think maybe he doesn't understand it like that. But it's, uh, Andreas Mokl. He's like a legend in OSS development. Check him out on GitHub. He has many projects. He's in, uh, software development for already decades, I guess. And he was, uh, or is active in Beehive community and they have relations to weather data because they like to measure temperature because bees tend to leave the hive when it gets cooler and they come back when it's warmer. Um, so there was his interest for weather data and also for the library. Yeah. And he joined me, helped me set up a clean structure for WetterDienst in the beginning and also some state-of-the-art CI-CD pipelines. And he also comes back every now and then. And then also, honestly, it pushes me also just to, uh, put, uh, again, some more time into WetterDienst, which is not always the case in the last month. And, uh, yeah. Uh, and then also, which is really important is open data because WetterDienst, uh, could not afford providing data when, like, for example, I would have to pay for the data. I would be bankrupt. Uh, but luckily there's open data these days. And I think it's great because it's, uh, enables research privately as well as institutionally. So, um, and this is great because in times of climate insecurity, we, it's good to have a single source of truth. So we can just look into the data and prove, um, the climate change we see outside maybe or in the models, which is also important, but also in terms of global insecurity. So it's good to have multiple institutions because, you know, these days some institutions may, um, yeah, quit or not provide data anymore and run models. So it's really important to have, I think, multiple institutions like DVD and yeah. Um, also, yeah, like I said, many national weather services publish their data already, like WetterDienst, like Deutsche WetterDienst since 2017. And they do it really accurately. So basically I think everything they measure, there's a data set for, um, it's, it's a real, um, it's a real lot of, lots of, uh, data they provide. Um, but there's also other services, for example, in Europe, the Geosphere, um, it's basically the environmental agency of Austria. They also have, um, an API these days and NOAA, the, the US version of the environmental agency also publishes the data in 2018. Um, and they, they have the benefit, like in terms of, uh, the Delta data wealthy that they also, um, ingest some of the data off, um, the German weather service, but not all of it because they have, uh, tie, they are tied to fixed resolutions. They have data in daily resolution and also an hourly, but, uh, nothing more than that. Um, and also I forgot, uh, open source software is obviously also really important because WetterDienst relies on Polars, uh, FastAPI, uh, Click and many more. And, um, those are obviously also really important for WetterDienst to work. Um, yeah, now let's look at WetterDienst itself. So, the statement is retrieve the entire climatological history of a place or a location like, for example, say Darmstadt. And this in 10, less than 10 lines of code. And this is actually, um, uh, what's the case these days? So we just need to import, um, uh, like, um, combination of provider and network. So provider would be DVD in this case and the network, um, I call it observation. So it's observational historical data. It's not forecast data, but it's historical. Uh, and then you just define the parameters you want. For example, in this case, it's, uh, in daily resolution, the climate summary data set. And from that, the parameter temperature, I mean two meters. So this is a regular measurement and two meters height. And this, I think it's called English hut or something. They measure it in. Plus maybe some, uh, start and end date. Plus maybe some settings for the unit you want to have the temperature values in. Uh, yeah. And then as you see also, you need some kind of station. Uh, no, first, first the results. And the, yeah, like I said, the results, the core results are, uh, Polar's data frames. I think everyone can work with it, uh, with it. Even like in the university, everybody knows data frames. Everybody loves them. Um, and the, on the top you see first, um, the station list. So basically you see on top we filter for station Dresden Klotsche. And then you, you get one row with all the information for that station. So it's basically, again, the, the parameter is listed there. Like daily and climate summary. Then the ID. Then, um, start and end date. And some geographical, um, information which is important to filter it out in a way. Um, and then the actual values. So there you get, uh, the list of values. Um, you also have a quality column on the, in the back, but we won't talk about that today. Um, yeah. And here you also see the, the, the details about the request. So it's daily climate summary and then the parameters. So the, the response is full of the details we, we have or we need from the request. Um, then there's also a metadata model, um, which is important to define all kinds of no-code stuff to, um, that, that basically describes the data set and its hierarchy. So there's some information about the provider and the network, some kind of metadata. Then there's a resolution which like, uh, shows how the values, the resolution of the values that you retrieve basically. So daily or hourly or whatever. The data set. So it's, it's data, the data sets name and, um, uh, some other definitions. And then maybe most importantly, the parameter model where you also have this definition of unit.

Speaker 2 [14:20]

Sorry.

Speaker 1 [14:21]

And, uh, the benefit of it is also that you can rewrite your request using this model with a dot annotation. Yeah. Ah, and it's, it's based on pedantic. I forgot to mention. So, um, and the other thing you need is the simplest form. You would just have to define two, um, or to do some implementations for, um, two abstract methods. Which one is, uh, one is, uh, all method for the station list. And the other one is, uh, collect station parameter data set. And this second method does all the collection for the data itself. So it fetches some data from the API. Um, there's different methods for the selection of the station. So typically in the beginning, you don't know which station you actually are looking for. So who knows the ID for the Darmstadt station? I don't know. Uh, so in the beginning, you could just request, uh, request all the stations. And then from this list on, um, um, yeah, look, look up the ID of the Darmstadt station. And once you have that, you could just go with, um, filter by station ID. Or, um, alternatively, you could do a fuzzy search for Darmstadt as a city name. Or you do some bounding box search, um, if you are interested in stations in a certain bounding box. You could also do a ranked search for, um, X number of stations closest to latitude-longitude pair. Um, or just retrieve all the stations in certain distance. But there's also, um, SQL filtering for the stations you want to retrieve data for. Um, yeah. And if no station is found, there's also interpolation and summary methods these days. And they try to automatically fetch as many stations as reasonable. So, um, you don't typically want to fetch all, like, 100 stations that would be able to use for interpolation. Because it takes ages. So instead we do, like, uh, I think, um, we try to fetch as many stations as feasible to interpolate, uh, as much data as possible in a certain date range. And this is important to notice. This works best for homogenous data, like temperature. Temperature is really equal over the area. But precipitation is really inhomogeneous over the area. So it's really different if you're looking at precipitation in Darmstadt or in Frankfurt, for example. Um, there's also many different exports. So from the data frame you can go on to some databases. You have file exports. Uh, and you have also many different, uh, formats also in Python objects. Uh, so you can just continue with, um, um, your, your, your, whatever, whatever setup of database you have locally. Um, there's also a CLI. I think everybody has a CLI today and also a Vetadienst. So this thing also replicates a request you've seen before. Um, yeah. And, um, you could just write, rewrite the, the request and with the CLI, but also there's, uh, if you install fast API optionally, you can do the same with an HTTP request. And the same, I show you the same for one for stations and the other thing for the, for the values. Um, yeah. Um, yeah. And there's also app these days. Um, shame on me. I've got good at it because I'm not a front end developer. Um, and the idea here is just to simplify the access for people who cannot code or set up a VN for whatever. Like my, I think my mother, for example, and the idea is, yeah, you have different, um, applications. One is the Explorer where you can just also, um, replicate the requests I've just, I've just shown you. Uh, just in a UI way. There's also climate stripes, stripes. I, uh, you may have seen them already some, on some events. It's a nice way to express climate change for your city. Um, yeah. Maybe I can also show it later to you if you want. Um, station history is, um, showing the metadata of some stations, but this metadata is for many services not available. I think at this time only for DVD. And there's also the REST API itself. So basically the HTTP request I've just shown you can just also send it to this, to this app, which is running on betterdienst. Eops.org. Uh, yeah. One more thing. This is a betterdienst, uh, the, this is a DVD API. So it's a plain file server. Uh, there's no actual API. And this is also the, the, the best reason why betterdienst actually started because you would just have to go all, through all the files manually and look it up yourself. You see even here the station list is not a real CSV. You have to do some manual parsing there. I like fuzzy parsing. Uh, yeah. But I think we are still lucky because they provide just so much data, uh, DVD. For example, this, uh, extensive dataset description where you find all the information about the parameter. And there are also some quirks, but I think we don't have too much time for this. Sorry to skip that. Yeah. And then the, the, what value did it give us? So I think for me it's fun. And also obviously the result belongs to me and you and everyone. And it's not hidden behind a business case. Um, it's a good place to try out new stuff. Um, and then usually, um, if it works like new tooling, I also tend to take it to my workplace and implement it there. Um, but there's also many, uh, impact on many businesses and processes. For example, agriculture, construction industry, for example, where they rely on temperature forecasts for concrete drying. Um, insurances, for example, where they want to show that a hazardous event was caused by this historical rainfall, for example. Uh, autonomous driving I've also seen. So people use a high resolution precipitation data for training on the rain sensors. Um, and then obviously also the research on climate change itself. Um, but my hope, um, I think, yeah, people who, who are into this already know how to get the data or they are in reach with, with me. Um, but my hope is also for you to inform yourself. So, but I'm still figuring out how to, how to improve Vetadienst or the Vetadienst app for this case. Yeah. Now the demo and, uh, Vibe code alert again. Uh, what I did is because I, I think I, what I wouldn't, I wouldn't want to rely on the internet here. So I just dumped the daily climate summary data set into a DuckDB file. Um, and I created a Marimo notebook around it. The DuckDB file I put also online. So if you want, you can just do your own SQL on it. So I jumped over to PyCharm. Just started. And I did some analysis of, uh, some analysis of this data. So first thing, um, we see is, um, yeah, here's, here's, uh, by the way, the code snippet, how I created the DuckDB file. So I just use some to target function here to pump the data into DuckDB file. Yeah. And the first thing, first interesting thing we see is, uh, station network of this climate data, uh, climate summary data set. We see there are a total of 1,200 stations, but currently there's only 566. So there's been, obviously lots of changes going on in the past. We see also, uh, the historical development. Yeah. The first station was, um, opened in 1759. Um, yeah. And then we see by 1900, it was still quite scattered. 1950, it looks already better. Um, yeah. Um, yeah. And today we have a really dense, dense network of stations here. Um, what I asked, um, like, uh, what question I asked myself is, uh, impact of war. And you could see here that in the first world war, there was no real impact, but, uh, I think it was also not so many stations maybe. But then in world war two, you could definitely, uh, definitely see a decline in stations during the end. So I think, um, you always also see this kind of historic events set and sadly in this case, um, also in the data. So there's also some gaps in the data. Usually you see during the world war two period for Germany. Uh, yeah. And there's another interesting effect. I think we talked also doing my, uh, work at the Institute, uh, at the, at the university. And you see here in the, in the end, you usually would expect, well, we have developed a good network. Now maybe let's continue or even, or just maintain the same number of stations. But there's actually a decline here. And I think there were two reasons. One is that usually the network is, um, that there's volunteers working on it. Usually retired, uh, teachers, I, uh, I heard. And they have to go to the station every day and check if everything is okay. And actually I think this is hard to, um, yeah, acquire those people, those volunteers these days. Uh, and it comes also with a side effect. So DVD, as at least when we, when I was in the Institute, we were discussing this. So DVD, I think, tended to, um, replace, uh, people by, um, like automated technique, like cameras and this kind of stuff. And they were actually not really happy with it because it would, or they, they would implicate that the quality of the measurements would go down with this. Um, yeah. With this change. All right.

Speaker 2 [24:45]

Well, I guess we are short on time. Uh, we have five minutes for the QI session. If that's okay. If that's okay, I will, uh, read out the questions one by one. Okay. The first question. Um, is, uh, interpolation between multiple stations purely mathematical or does use some kind of meteorological model in the background?

Speaker 1 [25:15]

Um, yeah, good, good question because, um, also you have some, um, sometimes difference in the height of the station and this kind of stuff. So, yeah, geometry also, I think should play a role, but honestly, uh, at the moment, uh, I can say no, except we do some, uh, additional technique for precipitation interpolation. Um, which is because there are days, um, where, where there's zero precipitation, but usually when you interpolate, you always get not zero, but really close to it. So, um, for the case of interpolation of precipitation, we do another interpolation of has it rained or has it not rained. So zero, um, zero one, and then we do the cutoff there to say the interpolated value is either zero or above zero. Yeah, makes sense.

Speaker 2 [26:10]

Uh, the next question is, as Polars is relatively new, did you start with Pandas? Did you have any conversion problems or do you miss any features in Polars still?

Speaker 1 [26:23]

Um, yes, we started with Pandas, um, and, but like I said, we like to play with the new stuff and we were, um, yeah, happy to see this advertisement by Polars. It's really fast. It has lots of optimization. Honestly, I don't know what it means in reality, but we, uh, switch directly to Polars, um, when we thought it would have the most necessary features. And I think in the beginning we had also some workarounds. I think these days it's quite complete for what we need. So we do just some cleanup mostly on the data. And then, uh, yeah, for that case, it, it is already in state.

Speaker 2 [27:04]

Next question. The logic is set up so that if a good weather station isn't found, it pulls data from several nearby ones. But does it take elevation into account with just a straight line distance? Whether you actually get similar data seems like it would depend on the station. Um, can you repeat? Yes. Uh, the logic is set up, uh, so that if a good weather station isn't found, it pulls data from several nearby ones. But does it take elevation into account? No, for the elevation, no.

Speaker 1 [27:47]

At the moment, no. No.

Speaker 2 [27:49]

All right. Uh, what amounts of data do we talk about when working with weather data?

Speaker 1 [27:59]

Um, so I've shown you like on the, um, daily, um, so those daily values, obviously, I think are the ones that reach most back in time. Because it is just one time per day measurement. Uh, one person could go to the station, check the value and, like, return home. Um, and those values reach into, um, 18th century. So it's a really long time series. I think you find maybe only older ones in England. Um, and for the higher resolutions, they don't reach back that long. So for example, we have this one minute precipitation data. And this is starts, I think at, uh, in the nineties, 1990s. Um, yeah. And I think also, yeah. And that's, that's the answer. One minute and 10 minutes don't reach back that long, maybe in the nineties.

Speaker 2 [28:50]

I guess, uh, we have time for one more question. Which kind of distance are you using when searching for stations?

Speaker 1 [28:58]

Um, like the, the units maybe they mean? Yeah, right. Um, I don't, I don't find the name right now in my brain, but obviously we cannot just, um, we have to do some, like, uh, projection distance calculation. I don't know if that answers your question correctly, but I don't have the name in my brain right now.

Speaker 2 [29:21]

Thank you. You can ask him after the session. Well, thank you so much. Uh, okay. Let's thank our speaker once again. We'll come for a speaker once again. We'll come for a speaker once again. We'll come for a speaker once again. We'll come for a speaker once again. We'll come for a speaker once again.

Benjamin

Benjamin Gutzmann is a 32 year old Python/data engineer and maintainer of Wetterdienst, currently at Otto Group data.works (Data Engineer since 2023; previously Junior Data Engineer), working across Generative AI and data engineering on GCP with Python, SQL, Argo, and Terraform. He has built the Wetterdienst library at earth observations (hobby project, since 2018). Before his start into work life he has studied Hydrology (BSc, MSc) at TU Dresden.

Social card for talk: Wetterdienst: Fast, Unified Access to Open Weather Data with Polars