...

FONDA – Foundations of Workflows for Large-Scale Scientific Data Analysis

Besök webbplats

Om arbetsgivaren

DFG Collaborative Research Center 1404 in Berlin / Potsdam, Germany

“Human productivity arguably still is the most expensive resource, trumping power, performance, and other factors”

Essentially all scientific disciplines are generating an ever-increasing amount of data. To derive scientific discoveries, these data sets are analyzed by complex data analysis workflows (DAWs), which are series of discrete analysis programs arranged in (often non-linear) pipelines. Because they usually deal with very large data sets, DAWs must be executed on distributed and/or parallel computational infrastructures. Traditionally, DAWs are optimized for speed, which leads to solutions that are hard to reproduce and share and that are tightly bound to exactly one type of input. However, as stated as summary in a recent NSF/DOE workshop that brought together the workflow and the HPC communities, “… human productivity arguably still is the most expensive resource, trumping power, performance, and other factors …” [DOE15].

"Our long-term goal is to develop methods and tools that achieve substantial reductions in development time and development cost of Data Analysis Workflows"

The proposed CRC FONDA – “Foundations of workflows for large-scale scientific data analysis” – will take up this observation and investigate methods for increasing productivity in the development, execution, and maintenance of DAWs for large scientific data sets. Our long-term goal is to develop methods and tools that achieve substantial reductions in development time and development cost of DAWs. We will approach these questions from a fundamental perspective, i.e., we aim at finding new abstractions, models, and algorithms that can eventually form the basis of a new class of future DAW infrastructures. Toward these goals, FONDA in its first phase will focus on three critical properties of DAWs and of DAW engines, namely portability, adaptability, and dependability (PAD). We want to investigate answers to questions such as: How can we build DAWs and DAW engines that enable portability of analysis across different infrastructures? How must DAWs be designed to adapt to changing input data or slightly changing requirements? How can we build dependable DAW systems that are aware of and control their own limitations and preconditions?

"Data Analysis Workflows are bridges between two worlds"

DAWs are bridges between two worlds: First, the specific scientific discipline using a DAW, and, second, Computer Science, which builds the infrastructures necessary for developing and executing DAWs. Developing novel foundations for scientific DAWs thus requires a close interaction between these two worlds. FONDA implements this idea by building on an interdisciplinary group of PIs from Computer Science, Material Science, Geosciences, and the Life Sciences. Through these cooperations, FONDA’s research results will be continuously validated using relevant and current scientific problems from different fields of the natural sciences.

 [DOE15] DOE Workshop Report (2015): “The Future of Scientific Workflows – Report of the NGNS/DOE Scientific Workflows Workshop”  

Arbetsgivarplats

Liknande arbetsgivare

...
KTH Royal Institute of Technology Stockholm, Sverige 107 lediga jobb
...
University of Luxembourg Luxemburg 100 lediga jobb
...
KU Leuven Leuven, Belgien 94 lediga jobb
...
Ghent University Gent, Belgien 69 lediga jobb
...
University of Twente Enschede, Nederländerna 62 lediga jobb
Fler arbetsgivare

Intressanta artiklar

...
5 Reasons to Pursue Your PhD at EMBL European Molecular Biology Laboratory (EMBL) 4 min läsning
...
Deciphering the Gut’s Clues to Our Health University of Turku 5 min läsning
...
Understanding Users to Optimise 3D Experiences Centrum Wiskunde & Informatica (CWI) 5 min läsning
...
Harnessing the Rhizosphere to Protect Our Soil Free University of Bozen - Bolzano 5 min läsning
Fler stories