Skip to content

Commit efdbebb

Browse files
author
Felix Hennig
committed
Add descriptions and fix formatting
1 parent bffebdd commit efdbebb

14 files changed

Lines changed: 279 additions & 300 deletions

‎docs/modules/demos/pages/airflow-scheduled-job.adoc‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
= airflow-scheduled-job
22
:page-aliases: stable@stackablectl::demos/airflow-scheduled-job.adoc
3+
:description: This demo installs Airflow with Postgres and Redis on Kubernetes, showcasing DAG scheduling, job runs, and status verification via the Airflow UI.
34

45
Install this demo on an existing Kubernetes cluster:
56

@@ -102,9 +103,10 @@ Click on the `run_every_minute` box in the centre of the page and then select `L
102103

103104
[WARNING]
104105
====
105-
In this demo, the logs are not available when the KubernetesExecutor is deployed. See the https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/executor/kubernetes.html#managing-dags-and-logs[Airflow Documentation] for more details.
106+
In this demo, the logs are not available when the KubernetesExecutor is deployed.
107+
See the https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/executor/kubernetes.html#managing-dags-and-logs[Airflow Documentation] for more details.
106108
107-
If you are interested in persisting the logs, please take a look at the xref:logging.adoc[] demo.
109+
If you are interested in persisting the logs, take a look at the xref:logging.adoc[] demo.
108110
====
109111

110112
image::airflow-scheduled-job/airflow_9.png[]

‎docs/modules/demos/pages/data-lakehouse-iceberg-trino-spark.adoc‎

Lines changed: 41 additions & 42 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
= data-lakehouse-iceberg-trino-spark
22
:page-aliases: stable@stackablectl::demos/data-lakehouse-iceberg-trino-spark.adoc
3+
:description: This demo shows a data workload with real-world data volumes using Trino, Kafka, Spark, NiFi, Superset and OPA.
34

45
:demo-code: https://github.com/stackabletech/demos/blob/main/demos/data-lakehouse-iceberg-trino-spark/create-spark-ingestion-job.yaml
56
:iceberg-table-maintenance: https://iceberg.apache.org/docs/latest/spark-procedures/#metadata-management
@@ -11,17 +12,17 @@
1112

1213
[IMPORTANT]
1314
====
14-
This demo shows a data workload with real-world data volumes and uses significant resources to ensure acceptable
15-
response times. It will likely not run on your workstation.
15+
This demo shows a data workload with real-world data volumes and uses significant resources to ensure acceptable response times.
16+
It will likely not run on your workstation.
1617
17-
There is also the smaller xref:trino-iceberg.adoc[] demo focusing on the abilities a lakehouse using Apache
18-
Iceberg offers. The `trino-iceberg` demo has no streaming data and can be executed on a local workstation.
18+
There is also the smaller xref:trino-iceberg.adoc[] demo focusing on the abilities a lakehouse using Apache Iceberg offers.
19+
The `trino-iceberg` demo has no streaming data and can be executed on a local workstation.
1920
====
2021

2122
[CAUTION]
2223
====
23-
This demo only runs in the `default` namespace, as a `ServiceAccount` will be created. Additionally, we have to use the
24-
FQDN service names (including the namespace), so that the used TLS certificates are valid.
24+
This demo only runs in the `default` namespace, as a `ServiceAccount` will be created.
25+
Additionally, we have to use the FQDN service names (including the namespace), so that the used TLS certificates are valid.
2526
====
2627

2728
Install this demo on an existing Kubernetes cluster:
@@ -37,9 +38,9 @@ $ stackablectl demo install data-lakehouse-iceberg-trino-spark
3738
The demo was developed and tested on a kubernetes cluster with about 12 nodes (4 cores with hyperthreading/SMT, 20GiB RAM and 30GB HDD).
3839
Instance types that loosely correspond to this on the Hyperscalers are:
3940

40-
- *Google*: `e2-standard-8`
41-
- *Azure*: `Standard_D4_v2`
42-
- *AWS*: `m5.2xlarge`
41+
* *Google*: `e2-standard-8`
42+
* *Azure*: `Standard_D4_v2`
43+
* *AWS*: `m5.2xlarge`
4344
4445
In addition to these nodes the operators will request multiple persistent volumes with a total capacity of about 300Gi.
4546

@@ -49,26 +50,26 @@ This demo will
4950

5051
* Install the required Stackable operators.
5152
* Spin up the following data products:
52-
** *Trino*: A fast distributed SQL query engine for big data analytics that helps you explore your data universe. This
53-
demo uses it to enable SQL access to the data.
54-
** *Apache Spark*: A multi-language engine for executing data engineering, data science, and machine learning. This demo
55-
uses it to stream data from Kafka into the lakehouse.
56-
** *MinIO*: S3 compatible object store. This demo uses it as persistent storage to store all the data used
57-
** *Apache Kafka*: A distributed event streaming platform for high-performance data pipelines, streaming analytics and
58-
data integration. This demo uses it as an event streaming platform to stream the data in near real-time.
59-
** *Apache NiFi*: An easy-to-use, robust system to process and distribute data. This demo uses it to fetch multiple
60-
online real-time data sources and ingest it into Kafka.
61-
** *Apache Hive metastore*: A service that stores metadata related to Apache Hive and other services. This demo uses it
62-
as metadata storage for Trino and Spark.
63-
** *Open policy agent* (OPA): An open-source, general-purpose policy engine unifying policy enforcement across the
64-
stack. This demo uses it as the authorizer for Trino, which decides which user can query which data.
65-
** *Apache Superset*: A modern data exploration and visualization platform. This demo utilizes Superset to retrieve data
66-
from Trino via SQL queries and build dashboards on top of that data.
53+
** *Trino*: A fast distributed SQL query engine for big data analytics that helps you explore your data universe.
54+
This demo uses it to enable SQL access to the data.
55+
** *Apache Spark*: A multi-language engine for executing data engineering, data science, and machine learning.
56+
This demo uses it to stream data from Kafka into the lakehouse.
57+
** *MinIO*: S3 compatible object store.
58+
This demo uses it as persistent storage to store all the data used
59+
** *Apache Kafka*: A distributed event streaming platform for high-performance data pipelines, streaming analytics and data integration.
60+
This demo uses it as an event streaming platform to stream the data in near real-time.
61+
** *Apache NiFi*: An easy-to-use, robust system to process and distribute data.
62+
This demo uses it to fetch multiple online real-time data sources and ingest it into Kafka.
63+
** *Apache Hive metastore*: A service that stores metadata related to Apache Hive and other services.
64+
This demo uses it as metadata storage for Trino and Spark.
65+
** *Open policy agent* (OPA): An open-source, general-purpose policy engine unifying policy enforcement across the stack.
66+
This demo uses it as the authorizer for Trino, which decides which user can query which data.
67+
** *Apache Superset*: A modern data exploration and visualization platform.
68+
This demo utilizes Superset to retrieve data from Trino via SQL queries and build dashboards on top of that data.
6769
* Copy multiple data sources in CSV and Parquet format into the S3 staging area.
68-
* Let Trino copy the data from the staging area into the lakehouse area. During the copy, transformations such as
69-
validating, casting, parsing timestamps and enriching the data by joining lookup tables are done.
70-
* Simultaneously, start a NiFi workflow, which fetches datasets in real-time via the internet and ingests the data as
71-
JSON records into Kafka.
70+
* Let Trino copy the data from the staging area into the lakehouse area.
71+
During the copy, transformations such as validating, casting, parsing timestamps and enriching the data by joining lookup tables are done.
72+
* Simultaneously, start a NiFi workflow, which fetches datasets in real-time via the internet and ingests the data as JSON records into Kafka.
7273
* Spark structured streaming job is started, which streams the data out of Kafka into the lakehouse.
7374
* Create Superset dashboards for visualization of the different datasets.
7475

@@ -83,9 +84,8 @@ As Apache Iceberg states on their https://iceberg.apache.org/docs/latest/[websit
8384
> Apache Iceberg is an open table format for huge analytic datasets. Iceberg adds tables to compute engines including
8485
Spark, Trino, PrestoDB, Flink, Hive and Impala using a high-performance table format that works just like a SQL table.
8586

86-
This demo uses Iceberg, which plays nicely with object storage and has integrations for Trino and Spark. It also
87-
provides the following benefits among other things, instead of putting https://parquet.apache.org/[Apache Parquet] files
88-
directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hive connector]:
87+
This demo uses Iceberg, which plays nicely with object storage and has integrations for Trino and Spark.
88+
It also provides the following benefits among other things, instead of putting https://parquet.apache.org/[Apache Parquet] files directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hive connector]:
8989

9090
* *Standardized table storage:* Using this standardized specification, multiple tools such as Trino, Spark and Flink can
9191
read and write Iceberg tables.
@@ -101,11 +101,11 @@ directly into S3 using the https://trino.io/docs/current/connector/hive.html[Hiv
101101
column `day` is not needed anymore, and the query `select count(\*) where ts > now() - interval 1 day` would use
102102
partition pruning as expected to read only one the partitions from today and yesterday.
103103
* *Branching and tagging:* Iceberg enables git-like semantics on your lakehouse. You can create tags pointing to a
104-
specific snapshot of your data and branches. For details, please read
104+
specific snapshot of your data and branches. For details, read
105105
https://www.dremio.com/blog/exploring-branch-tags-in-apache-iceberg-using-spark/[this excellent blog post]. Currently,
106106
this is only supported in Spark. Trino is https://github.com/trinodb/trino/issues/12844[working on support].
107107

108-
If you want to read more about the motivation and the working principles of Iceberg, please have a read on their
108+
If you want to read more about the motivation and the working principles of Iceberg, have a read on their
109109
https://iceberg.apache.org[website] or https://github.com/apache/iceberg/[GitHub repository].
110110

111111
== List the deployed Stackable services
@@ -226,8 +226,7 @@ On the right side are three strands, that
226226
. Fetch the current shared bike station status
227227
. Fetch the current shared bike status
228228

229-
For details on the NiFi workflow ingesting water-level data, please read the
230-
xref:nifi-kafka-druid-water-level-data.adoc#_nifi[nifi-kafka-druid-water-level-data documentation on NiFi].
229+
For details on the NiFi workflow ingesting water-level data, read the xref:nifi-kafka-druid-water-level-data.adoc#_nifi[nifi-kafka-druid-water-level-data documentation on NiFi].
231230

232231
== Spark
233232

@@ -310,9 +309,10 @@ to strings and the json needs to be parsed.
310309
.withColumn("json", from_json(col("value"), schema)) \
311310
----
312311

313-
Afterwards, we only select the needed fields (coming from JSON). We are not interested in all the other fields, such as
314-
`key`, `value`, `topic` or `offset`. The metadata of the Kafka records, such as `topic`, `timestamp`, `partition` and
315-
`offset`, are also available. Please have a look at the {spark-streaming-docs}[Spark streaming documentation on Kafka].
312+
Afterwards, we only select the needed fields (coming from JSON).
313+
We are not interested in all the other fields, such as `key`, `value`, `topic` or `offset`.
314+
The metadata of the Kafka records, such as `topic`, `timestamp`, `partition` and `offset`, are also available.
315+
Have a look at the {spark-streaming-docs}[Spark streaming documentation on Kafka].
316316

317317
[source,python]
318318
----
@@ -459,8 +459,7 @@ data files causes an unnecessary amount of metadata and less efficient queries f
459459
data files in parallel using Spark with the rewriteDataFiles action. This will combine small files into larger files to
460460
reduce metadata overhead and runtime file open cost.
461461

462-
Some tables will also be sorted during rewrite, please have a look at the
463-
{iceberg-rewrite}[documentation on rewrite_data_files].
462+
Some tables will also be sorted during rewrite, have a look at the {iceberg-rewrite}[documentation on rewrite_data_files].
464463

465464
== Trino
466465

@@ -479,8 +478,8 @@ image::data-lakehouse-iceberg-trino-spark/trino_2.png[]
479478

480479
=== Connect to Trino
481480

482-
Please have a look at the xref:home:trino:usage-guide/connect_to_trino.adoc[trino-operator documentation on how to
483-
connect to Trino]. This demo recommends to use DBeaver, as Trino consists of many schemas and tables you can explore.
481+
Have a look at the xref:home:trino:usage-guide/connect_to_trino.adoc[trino-operator documentation on how to connect to Trino].
482+
This demo recommends to use DBeaver, as Trino consists of many schemas and tables you can explore.
484483

485484
image::data-lakehouse-iceberg-trino-spark/dbeaver_1.png[]
486485

‎docs/modules/demos/pages/end-to-end-security.adoc‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
= end-to-end-security
2-
32
:k8s-cpu: https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-metrics-pipeline/#cpu
3+
:description: This demo showcases end-to-end security in Stackable Data Platform with OPA, featuring row/column access control, OIDC, Kerberos, and flexible group policies.
44

55
This is a demo to showcase what can be done with Open Policy Agent around authorization in the Stackable Data Platform.
66
It covers the following aspects of security:
@@ -55,8 +55,7 @@ You can see the deployed products and their relationship in the following diagra
5555

5656
image::end-to-end-security/overview.png[Architectural overview]
5757

58-
Please note the different types of arrows used to connect the technologies in here, which symbolize
59-
how authentication happens along that route and if impersonation is used for queries executed.
58+
Note the different types of arrows used to connect the technologies in here, which symbolize how authentication happens along that route and if impersonation is used for queries executed.
6059

6160
The Trino schema (with schemas, tables and views) is shown below.
6261

‎docs/modules/demos/pages/hbase-hdfs-load-cycling-data.adoc‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
= hbase-hdfs-cycling-data
22
:page-aliases: stable@stackablectl::demos/hbase-hdfs-load-cycling-data.adoc
3+
:description: Load cyclist data from HDFS to HBase on Kubernetes using Stackable's demo. Install, copy data, create HFiles, and query efficiently.
34

45
:kaggle: https://www.kaggle.com/datasets/timgid/cyclistic-dataset-google-certificate-capstone?select=Divvy_Trips_2020_Q1.csv
56
:k8s-cpu: https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-metrics-pipeline/#cpu
@@ -195,7 +196,7 @@ COLUMN FAMILIES DESCRIPTION
195196
[TIP]
196197
====
197198
Run `stackablectl stacklet list` to get the address of the _ui-http_ endpoint.
198-
If the UI is unavailable, please do a port-forward `kubectl port-forward hbase-master-default-0 16010`.
199+
If the UI is unavailable, do a port-forward `kubectl port-forward hbase-master-default-0 16010`.
199200
====
200201

201202
The Hbase web UI will give you information on the status and metrics of your Hbase cluster. See below for the start page.
Lines changed: 17 additions & 20 deletions
Original file line numberDiff line numberDiff line change
@@ -1,33 +1,30 @@
11
= Demos
22
:page-aliases: stable@stackablectl::demos/index.adoc
3+
:description: Explore Stackable demos showcasing data platform architectures. Includes external components for evaluation.
34

4-
The pages below this section guide you on how to use the demos provided by Stackable. To install a demo please follow
5-
the xref:management:stackablectl:quickstart.adoc[quickstart guide] or have a look at the
6-
xref:management:stackablectl:commands/demo.adoc[demo command]. We currently offer the following list of demos:
5+
The pages in this section guide you on how to use the demos provided by Stackable.
6+
To install a demo follow the xref:management:stackablectl:quickstart.adoc[quickstart guide] or have a look at the xref:management:stackablectl:commands/demo.adoc[demo command].
7+
These are the available demos:
78

89
include::partial$demos.adoc[]
910

1011
[IMPORTANT]
1112
.External Components in these demos
1213
====
13-
These demos are provided by Stackable as showcases to demonstrate potential architectures that could be built with the
14-
Stackable Data Platform. As such they may include components that are not supported by Stackable as part of our
15-
commercial offering.
14+
These demos are provided by Stackable as showcases to demonstrate potential architectures that could be built with the Stackable Data Platform.
15+
As such they may include components that are not supported by Stackable as part of our commercial offering.
1616
17-
If you are evaluating one or more of these demos with the intention of purchasing a subscription, please make sure to
18-
double-check the list of supported operators, anything that is not mentioned on there is not part of our commercial
19-
offering.
17+
If you are evaluating one or more of these demos with the intention of purchasing a subscription, make sure to double-check the list of supported operators; anything that is not mentioned on there is not part of our commercial offering.
2018
21-
Below you can find a list of components that are currently contained in one or more of the demos for reference, if
22-
something is missing from this list and also not mentioned on our operators list, then this component is not supported:
19+
Below you can find a list of components that are currently contained in one or more of the demos for reference, if something is missing from this list and also not mentioned on our operators list, then this component is not supported:
2320
24-
- Grafana
25-
- JupyterHub
26-
- MinIO
27-
- OpenLDAP
28-
- OpenSearch
29-
- OpenSearch Dashboards
30-
- PostgreSQL
31-
- Prometheus
32-
- Redis
21+
* Grafana
22+
* JupyterHub
23+
* MinIO
24+
* OpenLDAP
25+
* OpenSearch
26+
* OpenSearch Dashboards
27+
* PostgreSQL
28+
* Prometheus
29+
* Redis
3330
====

0 commit comments

Comments
 (0)