Skip to content

Commit 6cddc91

Browse files
author
Felix Hennig
committed
Add descriptions and fix formatting
1 parent efdbebb commit 6cddc91

5 files changed

Lines changed: 71 additions & 73 deletions

File tree

‎docs/modules/demos/pages/data-lakehouse-iceberg-trino-spark.adoc‎

Lines changed: 61 additions & 63 deletions
Original file line numberDiff line numberDiff line change
@@ -230,13 +230,11 @@ For details on the NiFi workflow ingesting water-level data, read the xref:nifi-
230230

231231
== Spark
232232

233-
https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html[Spark Structured Streaming] is used to
234-
stream data from Kafka into the lakehouse.
233+
https://spark.apache.org/docs/latest/structured-streaming-programming-guide.html[Spark Structured Streaming] is used to stream data from Kafka into the lakehouse.
235234

236235
=== Accessing the web interface
237236

238-
To have access to the Spark web interface you need to run the following command to forward port 4040 to your local
239-
machine.
237+
To have access to the Spark web interface you need to run the following command to forward port 4040 to your local machine.
240238

241239
[source,console]
242240
----
@@ -249,23 +247,27 @@ image::data-lakehouse-iceberg-trino-spark/spark_1.png[]
249247

250248
=== Listing the running Structured Streaming jobs
251249

252-
The UI displays the last job runs. Each running Structured Streaming job creates lots of Spark jobs internally. Click on
253-
the `Structured Streaming` tab to see the running streaming jobs.
250+
The UI displays the last job runs.
251+
Each running Structured Streaming job creates lots of Spark jobs internally.
252+
Click on the `Structured Streaming` tab to see the running streaming jobs.
254253

255254
image::data-lakehouse-iceberg-trino-spark/spark_2.png[]
256255

257-
Five streaming jobs are currently running. You can also click on a streaming job to get more details. For the job
258-
`ingest smart_city shared_bikes_station_status` click, on the `Run ID` highlighted in blue to open them up.
256+
Five streaming jobs are currently running.
257+
You can also click on a streaming job to get more details.
258+
For the job `ingest smart_city shared_bikes_station_status` click, on the `Run ID` highlighted in blue to open them up.
259259

260260
image::data-lakehouse-iceberg-trino-spark/spark_3.png[]
261261

262262
=== How the Structured Streaming jobs work
263263

264-
The demo has started all the running streaming jobs. Look at the {demo-code}[demo code] to see the actual code
265-
submitted to Spark. This document will explain one specific ingestion job - `ingest water_level measurements`.
264+
The demo has started all the running streaming jobs. Look at the {demo-code}[demo code] to see the actual code submitted to Spark.
265+
This document will explain one specific ingestion job - `ingest water_level measurements`.
266266

267-
The streaming job is written in Python using `pyspark`. First off, the schema used to parse the JSON coming from Kafka
268-
is defined. Nested structures or arrays are supported as well. The schema differs from job to job.
267+
The streaming job is written in Python using `pyspark`.
268+
First off, the schema used to parse the JSON coming from Kafka is defined.
269+
Nested structures or arrays are supported as well.
270+
The schema differs from job to job.
269271

270272
[source,python]
271273
----
@@ -276,16 +278,13 @@ schema = StructType([ \
276278
])
277279
----
278280

279-
Afterwards, a streaming read from Kafka is started. It reads from our Kafka at `kafka:9093` with the topic
280-
`water_levels_measurements`. When starting up, the job will ready all the existing messages in Kafka (read from
281-
earliest) and will process 50000000 records as a maximum in a single batch. As Kafka has retention set up, Kafka records
282-
might alter out of the topic before Spark has read the records, which can be the case when the Spark application wasn't
283-
running or crashed for too long. In the case of this demo, the streaming job should not error out. For a production job,
284-
`failOnDataLoss` should be set to `true` so that missing data does not go unnoticed - and Kafka offsets need to be
285-
adjusted manually, as well as some post-loading of data.
281+
Afterwards, a streaming read from Kafka is started. It reads from our Kafka at `kafka:9093` with the topic `water_levels_measurements`.
282+
When starting up, the job will ready all the existing messages in Kafka (read from earliest) and will process 50000000 records as a maximum in a single batch.
283+
As Kafka has retention set up, Kafka records might alter out of the topic before Spark has read the records, which can be the case when the Spark application wasn't running or crashed for too long.
284+
In the case of this demo, the streaming job should not error out.
285+
For a production job, `failOnDataLoss` should be set to `true` so that missing data does not go unnoticed - and Kafka offsets need to be adjusted manually, as well as some post-loading of data.
286286

287-
*Note:* The following Python snippets belong to a single Python statement but are split into separate blocks for better
288-
explanation.
287+
*Note:* The following Python snippets belong to a single Python statement but are split into separate blocks for better explanation.
289288

290289
[source,python]
291290
----
@@ -300,8 +299,7 @@ spark \
300299
.load() \
301300
----
302301

303-
So far we have a `readStream` reading from Kafka. Records on Kafka are simply a byte-stream, so they must be converted
304-
to strings and the json needs to be parsed.
302+
So far we have a `readStream` reading from Kafka. Records on Kafka are simply a byte-stream, so they must be converted to strings and the json needs to be parsed.
305303

306304
[source,python]
307305
----
@@ -319,10 +317,10 @@ Have a look at the {spark-streaming-docs}[Spark streaming documentation on Kafka
319317
.select("json.station_uuid", "json.timestamp", "json.value") \
320318
----
321319

322-
After all these transformations, we need to specify the sink of the stream, in this case, the Iceberg lakehouse. We are
323-
writing in the `iceberg` format using the `update` mode rather than the "normal" `append` mode. Spark will aim for a
324-
micro-batch every 2 minutes and save its checkpoints (its current offsets on the Kafka topic) in the specified S3
325-
location. Afterwards, the streaming job will be started by calling `.start()`.
320+
After all these transformations, we need to specify the sink of the stream, in this case, the Iceberg lakehouse.
321+
We are writing in the `iceberg` format using the `update` mode rather than the "normal" `append` mode.
322+
Spark will aim for a micro-batch every 2 minutes and save its checkpoints (its current offsets on the Kafka topic) in the specified S3 location.
323+
Afterwards, the streaming job will be started by calling `.start()`.
326324

327325
[source,python]
328326
----
@@ -345,10 +343,9 @@ One important part was skipped during the walkthrough:
345343
.foreachBatch(upsertWaterLevelsMeasurements) \
346344
----
347345

348-
`upsertWaterLevelsMeasurements` is a Python function that describes inserting the records from Kafka into the lakehouse
349-
table. This specific streaming job removes all duplicate records that can occur because of how the PegelOnline API works
350-
and gets called. As we don't want duplicate rows in our lakehouse tables, we need to filter the duplicates out as
351-
follows.
346+
`upsertWaterLevelsMeasurements` is a Python function that describes inserting the records from Kafka into the lakehouse table.
347+
This specific streaming job removes all duplicate records that can occur because of how the PegelOnline API works and gets called.
348+
As we don't want duplicate rows in our lakehouse tables, we need to filter the duplicates out as follows.
352349

353350
[source,python]
354351
----
@@ -363,17 +360,16 @@ def upsertWaterLevelsMeasurements(microBatchOutputDF, batchId):
363360
""")
364361
----
365362

366-
First, the data frame containing the upserts (records from Kafka) will be registered as a temporary view so that they
367-
can be accessed via Spark SQL. Afterwards, the `MERGE INTO` statement adds the new records to the lakehouse table.
363+
First, the data frame containing the upserts (records from Kafka) will be registered as a temporary view so that they can be accessed via Spark SQL.
364+
Afterwards, the `MERGE INTO` statement adds the new records to the lakehouse table.
368365

369-
The incoming records are first de-duplicated (using `SELECT DISTINCT * FROM waterLevelsMeasurementsUpserts`) so that the
370-
data from Kafka does not contain duplicates. Afterwards, the - now duplication-free - records get added to the
371-
`lakehouse.water_levels.measurements`, but *only* if they still need to be present.
366+
The incoming records are first de-duplicated (using `SELECT DISTINCT * FROM waterLevelsMeasurementsUpserts`) so that the data from Kafka does not contain duplicates.
367+
Afterwards, the - now duplication-free - records get added to the `lakehouse.water_levels.measurements`, but *only* if they still need to be present.
372368

373369
=== The Upsert mechanism
374370

375-
The `MERGE INTO` statement can be used for de-duplicating data and updating existing rows in the lakehouse table. The
376-
`ingest water_level stations` streaming job uses the following `MERGE INTO` statement:
371+
The `MERGE INTO` statement can be used for de-duplicating data and updating existing rows in the lakehouse table.
372+
The `ingest water_level stations` streaming job uses the following `MERGE INTO` statement:
377373

378374
[source,sql]
379375
----
@@ -389,25 +385,25 @@ WHEN MATCHED THEN UPDATE SET *
389385
WHEN NOT MATCHED THEN INSERT *
390386
----
391387

392-
First, the data within a batch is de-deduplicated as well. The record containing the station update with the highest
393-
Kafka timestamp is the newest and will be used during Upsert.
388+
First, the data within a batch is de-deduplicated as well.
389+
The record containing the station update with the highest Kafka timestamp is the newest and will be used during Upsert.
394390

395-
If a record for a station (detected by the same `station_uuid`) already exists, its contents will be updated. If the
396-
station is yet to be discovered, it will be inserted. The `MERGE INTO` also supports updating subsets of fields and more
397-
complex calculations, e.g. incrementing a counter. For details, have a look at the
398-
{iceberg-merge-docs}[Iceberg MERGE INTO documentation].
391+
If a record for a station (detected by the same `station_uuid`) already exists, its contents will be updated.
392+
If the station is yet to be discovered, it will be inserted.
393+
The `MERGE INTO` also supports updating subsets of fields and more complex calculations, e.g. incrementing a counter.
394+
For details, have a look at the {iceberg-merge-docs}[Iceberg MERGE INTO documentation].
399395

400396
=== The Delete mechanism
401397

402-
The `MERGE INTO` statement can de-duplicate data and update existing lakehouse table rows. For details have a look at
403-
the {iceberg-merge-docs}[Iceberg MERGE INTO documentation].
398+
The `MERGE INTO` statement can de-duplicate data and update existing lakehouse table rows.
399+
For details have a look at the {iceberg-merge-docs}[Iceberg MERGE INTO documentation].
404400

405401
=== Table maintenance
406402

407403
As mentioned, Iceberg supports out-of-the-box {iceberg-table-maintenance}[table maintenance] such as compaction.
408404

409-
This demo executes some maintenance functions in a rudimentary Python loop with timeouts in between. When running in
410-
production, the maintenance can be scheduled using Kubernetes {k8s-cronjobs}[CronJobs] or {airflow}[Apache Airflow],
405+
This demo executes some maintenance functions in a rudimentary Python loop with timeouts in between.
406+
When running in production, the maintenance can be scheduled using Kubernetes {k8s-cronjobs}[CronJobs] or {airflow}[Apache Airflow],
411407
which the Stackable Data Platform also supports.
412408

413409
[source,python]
@@ -439,7 +435,8 @@ while True:
439435
time.sleep(25 * 60) # Assuming compaction takes 5 min run every 30 minutes
440436
----
441437

442-
The scripts have a dictionary of all the tables to run maintenance on. The following procedures are run:
438+
The scripts have a dictionary of all the tables to run maintenance on.
439+
The following procedures are run:
443440

444441
==== https://iceberg.apache.org/docs/latest/spark-procedures/#expire_snapshots[expire_snapshots]
445442

@@ -495,8 +492,8 @@ Here you can see all the available Trino catalogs.
495492

496493
== Superset
497494

498-
Superset provides the ability to execute SQL queries and build dashboards. Open the Superset endpoint
499-
`external-http` in your browser (http://87.106.122.58:32452 in this case).
495+
Superset provides the ability to execute SQL queries and build dashboards.
496+
Open the Superset endpoint `external-http` in your browser (http://87.106.122.58:32452 in this case).
500497

501498
image::data-lakehouse-iceberg-trino-spark/superset_1.png[]
502499

@@ -506,13 +503,13 @@ image::data-lakehouse-iceberg-trino-spark/superset_2.png[]
506503

507504
=== Viewing the Dashboard
508505

509-
The demo has created dashboards to visualize the different data sources. Select the `Dashboards` tab at the top to view
510-
these dashboards.
506+
The demo has created dashboards to visualize the different data sources.
507+
Select the `Dashboards` tab at the top to view these dashboards.
511508

512509
image::data-lakehouse-iceberg-trino-spark/superset_3.png[]
513510

514-
Click on the dashboard called `House sales`. It might take some time until the dashboards renders all the included
515-
charts.
511+
Click on the dashboard called `House sales`.
512+
It might take some time until the dashboards renders all the included charts.
516513

517514
image::data-lakehouse-iceberg-trino-spark/superset_4.png[]
518515

@@ -528,27 +525,28 @@ There are multiple other dashboards you can explore on your own.
528525

529526
=== Viewing Charts
530527

531-
The dashboards consist of multiple charts. To list the charts, select the `Charts` tab at the top.
528+
The dashboards consist of multiple charts.
529+
To list the charts, select the `Charts` tab at the top.
532530

533531
=== Executing arbitrary SQL statements
534532

535-
Within Superset, you can create dashboards and run arbitrary SQL statements. On the top click on the tab `SQL` ->
536-
`SQL Lab`.
533+
Within Superset, you can create dashboards and run arbitrary SQL statements.
534+
On the top click on the tab `SQL` -> `SQL Lab`.
537535

538536
image::data-lakehouse-iceberg-trino-spark/superset_7.png[]
539537

540-
On the left, select the database `Trino lakehouse`, the schema `house_sales`, and set `See table schema` to
541-
`house_sales`.
538+
On the left, select the database `Trino lakehouse`, the schema `house_sales`, and set `See table schema` to `house_sales`.
542539

543540
[IMPORTANT]
544541
====
545-
The older screenshot below shows how the table preview would look like. Currently, there is an https://github.com/apache/superset/issues/25307[open issue]
546-
with previewing trino tables using the Iceberg connector. This doesn't affect the execution the following execution of the SQL statement.
542+
The older screenshot below shows how the table preview would look like. Currently, there is an https://github.com/apache/superset/issues/25307[open issue] with previewing trino tables using the Iceberg connector.
543+
This doesn't affect the execution the following execution of the SQL statement.
547544
====
548545

549546
image::data-lakehouse-iceberg-trino-spark/superset_8.png[]
550547

551-
In the right textbox, you can enter the desired SQL statement. If you want to avoid making one up, use the following:
548+
In the right textbox, you can enter the desired SQL statement.
549+
If you want to avoid making one up, use the following:
552550

553551
[source,sql]
554552
----

‎docs/modules/demos/pages/hbase-hdfs-load-cycling-data.adoc‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -136,7 +136,7 @@ cycling-tripdata
136136
----
137137

138138
Secondly, we'll use `org.apache.hadoop.hbase.tool.LoadIncrementalHFiles` (see {bulkload}[bulk load docs]) to import
139-
the Hfiles into the table and ingest rows.
139+
the Hfiles into the table and ingest rows.
140140

141141
Now we will see how many rows are in the `cycling-tripdata` table:
142142

@@ -209,8 +209,7 @@ image::hbase-hdfs-load-cycling-data/hbase-table-ui.png[]
209209

210210
== Accessing the HDFS web interface
211211

212-
You can also see HDFS details via a UI by running `stackablectl stacklet list` and following the link next to one of
213-
the namenodes.
212+
You can also see HDFS details via a UI by running `stackablectl stacklet list` and following the link next to one of the namenodes.
214213

215214
Below you will see the overview of your HDFS cluster.
216215

‎docs/modules/demos/pages/nifi-kafka-druid-water-level-data.adoc‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -529,7 +529,7 @@ What might also be interesting is the average and current measurement of the sta
529529

530530
[source,sql]
531531
----
532-
select
532+
select
533533
stations.longname as station,
534534
avg("value") as avg_measurement,
535535
latest_by("value", measurements."__time") as current_measurement,

‎docs/modules/demos/pages/signal-processing.adoc‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -40,10 +40,10 @@ image::signal-processing/overview.png[]
4040

4141
== Data ingestion
4242

43-
The data used in this demo is a set of gas sensor measurements*.
44-
The dataset provides resistance values (r-values hereafter) for each of 14 gas sensors.
45-
In order to simulate near-real-time ingestion of this data, it is downloaded and batch-inserted into a Timescale table.
46-
It's then updated in-place retaining the same time offsets but shifting the timestamps such that the notebook code can "move through" the data using windows as if it were being streamed.
43+
The data used in this demo is a set of gas sensor measurements*.
44+
The dataset provides resistance values (r-values hereafter) for each of 14 gas sensors.
45+
In order to simulate near-real-time ingestion of this data, it is downloaded and batch-inserted into a Timescale table.
46+
It's then updated in-place retaining the same time offsets but shifting the timestamps such that the notebook code can "move through" the data using windows as if it were being streamed.
4747
The Nifi flow that does this can easily be extended to process other sources of (actually streamed) data.
4848

4949
== JupyterHub

‎stacks/stacks-v2.yaml‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,10 +1,11 @@
1+
---
12
stacks:
23
monitoring:
34
description: Stack containing Prometheus and Grafana
45
stackableRelease: 24.7
56
stackableOperators:
6-
- commons
7-
- listener
7+
- commons
8+
- listener
89
labels:
910
- monitoring
1011
- prometheus

0 commit comments

Comments
 (0)