You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit efdbebb
Browse filesBrowse the repository at this point in the historyBrowse files
:description: This demo installs Airflow with Postgres and Redis on Kubernetes, showcasing DAG scheduling, job runs, and status verification via the Airflow UI.
3
4
4
5
Install this demo on an existing Kubernetes cluster:
5
6
@@ -102,9 +103,10 @@ Click on the `run_every_minute` box in the centre of the page and then select `L
102
103
103
104
[WARNING]
104
105
====
105
-
In this demo, the logs are not available when the KubernetesExecutor is deployed. See the https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/executor/kubernetes.html#managing-dags-and-logs[Airflow Documentation] for more details.
106
+
In this demo, the logs are not available when the KubernetesExecutor is deployed.
107
+
See the https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/executor/kubernetes.html#managing-dags-and-logs[Airflow Documentation] for more details.
106
108
107
-
If you are interested in persisting the logs, please take a look at the xref:logging.adoc[] demo.
109
+
If you are interested in persisting the logs, take a look at the xref:logging.adoc[] demo.
The demo was developed and tested on a kubernetes cluster with about 12 nodes (4 cores with hyperthreading/SMT, 20GiB RAM and 30GB HDD).
38
39
Instance types that loosely correspond to this on the Hyperscalers are:
39
40
40
-
-*Google*: `e2-standard-8`
41
-
-*Azure*: `Standard_D4_v2`
42
-
-*AWS*: `m5.2xlarge`
41
+
**Google*: `e2-standard-8`
42
+
**Azure*: `Standard_D4_v2`
43
+
**AWS*: `m5.2xlarge`
43
44
44
45
In addition to these nodes the operators will request multiple persistent volumes with a total capacity of about 300Gi.
45
46
@@ -49,26 +50,26 @@ This demo will
49
50
50
51
* Install the required Stackable operators.
51
52
* Spin up the following data products:
52
-
** *Trino*: A fast distributed SQL query engine for big data analytics that helps you explore your data universe. This
53
-
demo uses it to enable SQL access to the data.
54
-
** *Apache Spark*: A multi-language engine for executing data engineering, data science, and machine learning. This demo
55
-
uses it to stream data from Kafka into the lakehouse.
56
-
** *MinIO*: S3 compatible object store. This demo uses it as persistent storage to store all the data used
57
-
** *Apache Kafka*: A distributed event streaming platform for high-performance data pipelines, streaming analytics and
58
-
data integration. This demo uses it as an event streaming platform to stream the data in near real-time.
59
-
** *Apache NiFi*: An easy-to-use, robust system to process and distribute data. This demo uses it to fetch multiple
60
-
online real-time data sources and ingest it into Kafka.
61
-
** *Apache Hive metastore*: A service that stores metadata related to Apache Hive and other services. This demo uses it
62
-
as metadata storage for Trino and Spark.
63
-
** *Open policy agent* (OPA): An open-source, general-purpose policy engine unifying policy enforcement across the
64
-
stack. This demo uses it as the authorizer for Trino, which decides which user can query which data.
65
-
** *Apache Superset*: A modern data exploration and visualization platform. This demo utilizes Superset to retrieve data
66
-
from Trino via SQL queries and build dashboards on top of that data.
53
+
** *Trino*: A fast distributed SQL query engine for big data analytics that helps you explore your data universe.
54
+
This demo uses it to enable SQL access to the data.
55
+
** *Apache Spark*: A multi-language engine for executing data engineering, data science, and machine learning.
56
+
This demo uses it to stream data from Kafka into the lakehouse.
57
+
** *MinIO*: S3 compatible object store.
58
+
This demo uses it as persistent storage to store all the data used
59
+
** *Apache Kafka*: A distributed event streaming platform for high-performance data pipelines, streaming analytics and data integration.
60
+
This demo uses it as an event streaming platform to stream the data in near real-time.
61
+
** *Apache NiFi*: An easy-to-use, robust system to process and distribute data.
62
+
This demo uses it to fetch multiple online real-time data sources and ingest it into Kafka.
63
+
** *Apache Hive metastore*: A service that stores metadata related to Apache Hive and other services.
64
+
This demo uses it as metadata storage for Trino and Spark.
65
+
** *Open policy agent* (OPA): An open-source, general-purpose policy engine unifying policy enforcement across the stack.
66
+
This demo uses it as the authorizer for Trino, which decides which user can query which data.
67
+
** *Apache Superset*: A modern data exploration and visualization platform.
68
+
This demo utilizes Superset to retrieve data from Trino via SQL queries and build dashboards on top of that data.
67
69
* Copy multiple data sources in CSV and Parquet format into the S3 staging area.
68
-
* Let Trino copy the data from the staging area into the lakehouse area. During the copy, transformations such as
69
-
validating, casting, parsing timestamps and enriching the data by joining lookup tables are done.
70
-
* Simultaneously, start a NiFi workflow, which fetches datasets in real-time via the internet and ingests the data as
71
-
JSON records into Kafka.
70
+
* Let Trino copy the data from the staging area into the lakehouse area.
71
+
During the copy, transformations such as validating, casting, parsing timestamps and enriching the data by joining lookup tables are done.
72
+
* Simultaneously, start a NiFi workflow, which fetches datasets in real-time via the internet and ingests the data as JSON records into Kafka.
72
73
* Spark structured streaming job is started, which streams the data out of Kafka into the lakehouse.
73
74
* Create Superset dashboards for visualization of the different datasets.
74
75
@@ -83,9 +84,8 @@ As Apache Iceberg states on their https://iceberg.apache.org/docs/latest/[websit
83
84
> Apache Iceberg is an open table format for huge analytic datasets. Iceberg adds tables to compute engines including
84
85
Spark, Trino, PrestoDB, Flink, Hive and Impala using a high-performance table format that works just like a SQL table.
85
86
86
-
This demo uses Iceberg, which plays nicely with object storage and has integrations for Trino and Spark. It also
87
-
provides the following benefits among other things, instead of putting https://parquet.apache.org/[Apache Parquet] files
88
-
directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hive connector]:
87
+
This demo uses Iceberg, which plays nicely with object storage and has integrations for Trino and Spark.
88
+
It also provides the following benefits among other things, instead of putting https://parquet.apache.org/[Apache Parquet] files directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hive connector]:
89
89
90
90
* *Standardized table storage:* Using this standardized specification, multiple tools such as Trino, Spark and Flink can
91
91
read and write Iceberg tables.
@@ -101,11 +101,11 @@ directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hiv
101
101
column `day` is not needed anymore, and the query `select count(\*) where ts > now() - interval 1 day` would use
102
102
partition pruning as expected to read only one the partitions from today and yesterday.
103
103
* *Branching and tagging:* Iceberg enables git-like semantics on your lakehouse. You can create tags pointing to a
104
-
specific snapshot of your data and branches. For details, please read
104
+
specific snapshot of your data and branches. For details, read
105
105
https://www.dremio.com/blog/exploring-branch-tags-in-apache-iceberg-using-spark/[this excellent blog post]. Currently,
106
106
this is only supported in Spark. Trino is https://github.com/trinodb/trino/issues/12844[working on support].
107
107
108
-
If you want to read more about the motivation and the working principles of Iceberg, please have a read on their
108
+
If you want to read more about the motivation and the working principles of Iceberg, have a read on their
109
109
https://iceberg.apache.org[website] or https://github.com/apache/iceberg/[GitHub repository].
110
110
111
111
== List the deployed Stackable services
@@ -226,8 +226,7 @@ On the right side are three strands, that
226
226
. Fetch the current shared bike station status
227
227
. Fetch the current shared bike status
228
228
229
-
For details on the NiFi workflow ingesting water-level data, please read the
230
-
xref:nifi-kafka-druid-water-level-data.adoc#_nifi[nifi-kafka-druid-water-level-data documentation on NiFi].
229
+
For details on the NiFi workflow ingesting water-level data, read the xref:nifi-kafka-druid-water-level-data.adoc#_nifi[nifi-kafka-druid-water-level-data documentation on NiFi].
231
230
232
231
== Spark
233
232
@@ -310,9 +309,10 @@ to strings and the json needs to be parsed.
:description: This demo showcases end-to-end security in Stackable Data Platform with OPA, featuring row/column access control, OIDC, Kerberos, and flexible group policies.
4
4
5
5
This is a demo to showcase what can be done with Open Policy Agent around authorization in the Stackable Data Platform.
6
6
It covers the following aspects of security:
@@ -55,8 +55,7 @@ You can see the deployed products and their relationship in the following diagra
Please note the different types of arrows used to connect the technologies in here, which symbolize
59
-
how authentication happens along that route and if impersonation is used for queries executed.
58
+
Note the different types of arrows used to connect the technologies in here, which symbolize how authentication happens along that route and if impersonation is used for queries executed.
60
59
61
60
The Trino schema (with schemas, tables and views) is shown below.
:description: Explore Stackable demos showcasing data platform architectures. Includes external components for evaluation.
3
4
4
-
The pages below this section guide you on how to use the demos provided by Stackable. To install a demo please follow
5
-
the xref:management:stackablectl:quickstart.adoc[quickstart guide] or have a look at the
6
-
xref:management:stackablectl:commands/demo.adoc[demo command]. We currently offer the following list of demos:
5
+
The pages in this section guide you on how to use the demos provided by Stackable.
6
+
To install a demo follow the xref:management:stackablectl:quickstart.adoc[quickstart guide] or have a look at the xref:management:stackablectl:commands/demo.adoc[demo command].
7
+
These are the available demos:
7
8
8
9
include::partial$demos.adoc[]
9
10
10
11
[IMPORTANT]
11
12
.External Components in these demos
12
13
====
13
-
These demos are provided by Stackable as showcases to demonstrate potential architectures that could be built with the
14
-
Stackable Data Platform. As such they may include components that are not supported by Stackable as part of our
15
-
commercial offering.
14
+
These demos are provided by Stackable as showcases to demonstrate potential architectures that could be built with the Stackable Data Platform.
15
+
As such they may include components that are not supported by Stackable as part of our commercial offering.
16
16
17
-
If you are evaluating one or more of these demos with the intention of purchasing a subscription, please make sure to
18
-
double-check the list of supported operators, anything that is not mentioned on there is not part of our commercial
19
-
offering.
17
+
If you are evaluating one or more of these demos with the intention of purchasing a subscription, make sure to double-check the list of supported operators; anything that is not mentioned on there is not part of our commercial offering.
20
18
21
-
Below you can find a list of components that are currently contained in one or more of the demos for reference, if
22
-
something is missing from this list and also not mentioned on our operators list, then this component is not supported:
19
+
Below you can find a list of components that are currently contained in one or more of the demos for reference, if something is missing from this list and also not mentioned on our operators list, then this component is not supported:
0 commit comments