An interactive, exportable reference architecture for the Databricks Data Intelligence Platform. No build step and no backend: a static folder that lazy-loads its industry boards, translations and reference data as plain files, translates into sixteen languages, and exports to PDF, PowerPoint, PNG, GIF, a standalone HTML copy, and editable files for draw.io, Lucidchart, Visio, Miro, Figma and Excalidraw. Serve it locally, or deploy it into your own Databricks workspace as an app your teams reach from the workspace navigation.
Most reference architectures are a picture. Someone drew it in a diagramming tool, exported a PNG, and pasted it into a deck. It is accurate on the day it is made, it cannot be interrogated, and adapting it to a specific customer means starting again in the source tool, which nobody has.
Databricks Reference Architecture is the same architecture as a live document:
- Every box is a real product, and clicking it opens what it is, what a customer would not learn from the name, its release stage, the documentation for the cloud you are on, the product page, and related boxes you can jump to.
- The cloud provider is a switch. Azure, AWS and GCP swap the storage, compute, identity and ingestion services, and every documentation link re-points at that cloud's own docs.
- The platform is drawn in five shapes, so the same architecture fits a 16:9 slide, a portrait page, or a layout that has to leave room for a third party in the middle.
- Release stage is a filter. Show GA only for a procurement conversation, or add Beta and the previews for a roadmap one.
- Industry is a switch too. Sixty-three industries, each one specialising the sources, ingestion, teams, apps, use cases and consumers, with the medallion layers pointing at that industry's own data model. Every use case and every team is a story you can open: click a use case for the problem it solves, who benefits, how it is built, the components it uses and the customer stories that prove it; click a team for its sub-personas and the use cases they care about.
- The flows are alive. Every connector is a solid arrow with a glowing dot gliding from source to target, each on its own timing, so the board reads as many independent data flows rather than one static picture. The motion is captured in the GIF export too.
- A guided tour walks a first-time viewer through every control and every zone, and opens the detail panel on a real box so they see what a click gives them. It runs itself once on a first visit and replays any time from the ◎ button in the toolbar.
- It speaks sixteen languages. A language switch translates the board content into any of sixteen languages, right-to-left for Arabic and Hebrew, while the toolbar stays in English and product and brand names are never translated. Each language is fetched only when it is picked.
- Every board has a share card. A share button posts the board straight to LinkedIn, X, Reddit and more, and every industry has its own Open Graph preview image, so a pasted link unfurls in Slack, Teams or a social feed with a picture and title tailored to that industry rather than a bare URL.
- It exports, and the export points back. PDF and PowerPoint open on a cover carrying the title and a link to the exact live board the file was made from, then an index of the sections, the architecture, and one detailed page or slide per item across four sections: Use Cases, Genie Agents, AI/BI Dashboards and Databricks Apps, each carrying the same detail the on-screen drawer shows, and a closing slide that links back to the platform and the live board. The deck follows the theme and the Branded/Categorized choice on screen, so a categorized dark board exports a categorized dark deck. Plus PNG, an animated GIF that keeps the flow moving, and a standalone HTML copy.
Most people meet Databricks one feature at a time: a notebook here, a job there, Unity Catalog in a governance review. This board puts the entire Data Intelligence Platform in one frame: the sources on the left; ingestion, the medallion layers, unified governance and the agentic layers through the middle; and the consumers, the teams and the cloud services around the outside, with the data flows drawn as live arrows so you can see how one part feeds the next. Every box is a real product, not a label, so clicking one tells you what it is, the thing the name does not give away, what it actually does, and the boxes it connects to. In one screen a newcomer gets the shape of the whole platform, and an architect gets the connections between its parts.
It is not a generic diagram either, and that is the point. The cloud switch redraws the storage, compute, identity and ingestion services as Azure, AWS or GCP and re-points every documentation link at that cloud, so you are looking at the platform on the cloud you actually run. The industry switch respecialises the sources, teams, use cases and consumers for your sector and points the medallion layers at that industry's own data model. And when you need it tied to one specific estate, edit mode, the YAML export and import, and the AI assistant let you retarget the board to a particular customer or use case and export it, as a PDF, a PowerPoint or a standalone page, that shows exactly how Databricks fits into that architecture. It is the difference between a picture of the platform and a picture of your platform.
Because every box carries its own facts and links, the board doubles as a way to
learn Databricks rather than only to show it. Open any product and the Learn
more section takes you to the cloud-specific documentation for the cloud you are
on (the Azure Databricks docs on Azure, docs.databricks.com on AWS and GCP),
the product page, and a blog or deep dive, so each box is a launch point into the
official material rather than a dead end. The release-stage filter teaches what is
generally available today against what is still in preview or beta. The use cases
spell out the problem each one solves, who benefits, how it is built and which
components it touches, with links to real customer stories, so you see how the
features are used and not only what they are named. The Genie Agents and AI/BI
dashboards show the question-and-answer layer that sits on top. There is a
Databricks YouTube channel section with curated videos verified against the
official channel for when you would rather watch than read, and the whole board
translates into sixteen languages, so a team can learn in its own. A new joiner, a
partner or a customer can explore the platform at their own pace, from the first
box to the deepest link, without a slide deck in between.
Sources on the left, the platform through the middle, consumers on the right, with the teams who use it and the cloud it runs on wrapped around the outside.
| Zone | What sits there |
|---|---|
| Sources | Structured, semi-structured, unstructured, streaming and IoT, external and partner data, and federation sources reached without copying |
| Cloud and 3rd-party ingestion | The cloud's own ETL services, and third-party ELT and streaming brokers that land data alongside the platform's native ingestion |
| Platform | Ingest, Agentic Apps, Agentic Work, Unified Governance, Agentic Data, Open Infrastructure and the medallion layers, drawn as one outline traced by a moving ring |
| Teams | The business and technical teams the platform is built for. A tile is a team, never a job title, and the roles inside it and the surfaces they touch are in the side panel |
| Consumers | BI and productivity tools, MCP and APIs, published data products, partners and platforms, operational systems, four AI/BI Dashboards and four Databricks Apps, and the agent harnesses that arrive from outside |
| Use cases and Genie Agents | What the platform is used for, in the band above the platform: ten use cases and four Genie Agents on every board, each Genie Agent labelled with the domain it serves, industry-neutral by default |
| Cloud services and integrations | The account's own storage, compute, key vault, catalog, identity and observability services |
| Signal | Meaning |
|---|---|
| Solid arrow | Data moves along it, and a glowing dot glides from source to target in the direction it actually flows. Each arrow runs on its own timing, so the board reads as many independent flows rather than one synchronised pulse |
| Dashed zone outline | A grouping, not a boundary that data crosses |
| The ring around the platform | One continuous outline: the platform is one product, not a stack of separate ones |
| Colour | Identifies the zone, never the status. Every zone keeps its own hue in every palette and both themes |
Fourteen controls, left to right, sitting in the header above the diagram, plus platform zoom on the diagram itself.
| Control | What it does |
|---|---|
| Share feedback | Opens a feedback form in a new tab, for a comment on the reference architecture itself |
| Share | Posts the board on screen to LinkedIn, X, Reddit and more; the link points straight to the industry showing, and unfurls with that industry's own preview card |
| Industry | Sixty-three industries plus Standard Reference Architecture, searchable, every entry on one line. Specialises everything outside the platform: sources, ingestion, teams, apps, use cases and consumers. The platform itself and the cloud services band do not change |
| Display: Branded / Categorized | Shows every product by its commercial brand name (Branded), or rooted back to its generic category (Categorized, the default) for a room that knows the category but not the product, so SABRE reads as Airline Reservation System. The choice applies everywhere at once, the tiles, the detail drawers and every download, and it is remembered between visits |
| Cloud | Azure, AWS, GCP. Swaps the cloud services band, the cloud ETL tiles and the federation sources, and re-points every documentation link at that cloud's own docs, including the Microsoft Learn pages on Azure |
| Dark / Light | Follows the operating system by default, and remembers an explicit choice. Downloads follow whatever is on screen |
| Palette | Thirteen colour schemes in three groups |
| Style | Five platform shapes |
| Stage | Filters the platform box by release stage |
| Download | PDF, PowerPoint with an editable appendix, PNG, GIF, HTML, and editable draw.io, Visio, SVG and Excalidraw files. Its Architecture descriptor section exports the board on screen, the reference or any industry, as an editable YAML file, and imports one back to open your own version in a new tab |
| Details panel dock | Pins the detail drawer to the side so it stays open while you click from box to box, instead of overlaying the board each time |
| Language | Translates the board content into any of sixteen languages, right-to-left for Arabic and Hebrew. The toolbar and menus stay in English, and product and brand names are never translated |
| Tour (the ◎ at the end of the toolbar) | A guided walk-through that spotlights each control and each zone in turn, opens the detail panel on a real box so you see what a click gives you, and runs itself once on a first visit. Replayable any time |
| Zoom (on the platform heading, not the toolbar) | Zooms into the platform on its own: hides sources, consumers, the apps band and the cloud services. Click again to restore. Exports respect it, and the button itself never appears in one |
Everything that acts on a diagram goes inert on a tab that has no diagram yet, so the toolbar cannot be used against nothing. Dark/Light and Cloud stay live, because they are global and are the two things worth setting before a diagram exists.
Industry, cloud, shape, palette, theme and platform zoom are each a URL
parameter, so a board can be linked to in the state it was read in:
?industry=airlines&cloud=aws&shape=h90&pal=nordic&theme=light&platform=1. That
is the link the PDF and PowerPoint covers carry. The Branded/Categorized choice
is remembered locally rather than in the link, so a shared URL opens in the
reader's own preferred naming.
The industry list follows the Databricks Industry Data Models catalogue and extends it: airlines is the reference, and sixty-two more are authored to the same depth, spanning finance, healthcare and life sciences, public sector, manufacturing and energy, retail and consumer, media, technology, professional services and more. Picking one rewrites the four zones outside the platform, and the medallion layers start pointing at that industry's own folder in the repository.
Every industry is written to the same schema as the airlines reference: real, current vendor systems in the sources and ingestion, teams broken into sub-personas, four apps and ten use cases, and each use case carrying a problem statement, its beneficiary, how it is built, the architecture components it touches, and links to matching Databricks customer stories. Use cases with a story sort ahead of those without, so the board leads with proof.
An industry board is held to the reference board's height, which matters more
than it sounds: the board is scaled to fit, so one industry carrying a few more
tiles than the rest would render every label in that industry smaller. Teams and
ingestion each get a fixed pocket, eight tiles and three groups, and industry
content is written to that budget rather than allowed to grow past it.
tools/heightgate.py measures every industry against the reference in all five
shapes and fails on a board that is taller, a label that is clipped, or a label
that only fits by wrapping.
Each name opens the live board specialised for that industry: its own sources
and ingestion, its own teams, its own four apps and ten use cases, its own
consumers, and the medallion layers pointing at that industry's data model. Add
&cloud=aws or &cloud=gcp to any link to open it on that cloud; the default
is Azure.
Sixty-three industries in all, plus the industry-neutral Standard Reference Architecture that every board starts from.
| Shape | Why it exists |
|---|---|
| Z | The default. Ingest reaches left over the sources, serving reaches right over the consumers |
| S | The Z mirrored, for when the story runs right to left |
| T | Both arms on top, for a wide slide with a short middle |
| T180 | Both arms underneath |
| H90 | A full I-beam: three rows with split arms and pockets inside the notches, which is what leaves room for the cloud and third-party ingestion to sit inside the shape rather than beside it |
Every shape is measured, not eyeballed: the two pockets are equalised to the taller one, so the arms stay symmetrical to the pixel in all five.
H90 works differently enough from the other four to be worth spelling out. Each pocket splits into two labelled boxes side by side, Cloud ETL beside 3rd Party and Business beside Technical, and each box runs its tiles down a single column. Both pockets sit inside their own arm column rather than spanning the middle, so the crossbar of the I stays clear, and the governance band renders there instead of in the lower block: the narrow middle of the shape carries the platform's control plane rather than being a spacer. The arms are wider here than in the other shapes to fit those boxes, and the design canvas is wider to match, which is free because the fit in this shape is bound by height rather than width.
| Group | Palettes |
|---|---|
| Neutral | Spectrum (default), Mono (print safe), Muted (low chroma), Nordic (cool calm) |
| Coloured | Ocean (analogous), Earth (warm neutral), Sunset (warm shift), Berry (cool warm) |
| Loud | Solid (filled), Jewel (deep), Vivid (projector), Pop (playful), Neon (maximum) |
tools/palgen.py generates all of them. For every zone hue it walks lightness
until four values clear WCAG AA against the surface each one actually sits on,
in both themes: the zone fill, the chip tile inside it, the border, and the ink
on top. Nothing here is hand-picked, which is why Neon is legible and Mono
survives a black and white printer.
Spectrum keeps the chips white so the reference reads as a document. Every other palette tints the chips too, so a palette choice is visible in every box rather than in the labels alone.
| Stage | |
|---|---|
| GA | Available now |
| Public Preview | |
| Beta | |
| Private Preview | |
| Coming soon |
Each row carries the number of boxes it would show, and switching a stage off
removes those boxes from the platform, closes the gap and re-traces the ring.
Unstaged boxes always stay. An st value the app does not recognise is treated
as unstaged rather than quietly filed as available, because the conservative
reading is the honest one in front of a customer.
| Format | What you get |
|---|---|
| A cover with the title, the industry and cloud, the sentence the board leads with, and clickable links to the exact live board and to the industry's data model; an index page listing the sections; the architecture as a pixel-exact image of the board, in the current theme and palette; then four section breaks, each with one page per item, Use Cases (problem, beneficiary, build, components, story links), Genie Agents (the domain it serves, the data it reads, the teams and its top questions), AI/BI Dashboards (the KPIs it tracks and the teams that run on it) and Databricks Apps; then a closing page linking to the platform and the live board. All in the theme on screen and in Branded or Categorized names to match the board | |
| PowerPoint | The same pages, as native slides: a cover, an index slide, a board slide carrying a pixel-exact image of the board, then the Use Cases, Genie Agents, AI/BI Dashboards and Databricks Apps sections a slide per item, and a closing slide. An Editable Architecture appendix follows: a break slide, then the board again as grouped native shapes you can ungroup and edit in PowerPoint, Keynote or Google Slides. The deck follows the theme and the Branded/Categorized choice on screen, so a categorized dark board exports a categorized dark deck |
| PNG | 2x raster of the current view |
| GIF | A looping animation, 1400px wide, twelve frames, that keeps the travelling dashes and the platform ring moving. Roughly 200 KB, because only the moving pixels are stored per frame, in the palette and theme on screen |
| HTML | A standalone copy of the page with your current choices baked in, which opens anywhere with no server |
Edit in another app. These downloads rebuild the current view as grouped, editable shapes instead of a picture. Every zone, group box and tile is a group you can ungroup; labels are real text, arrows are connectors, product logos are crisp images and each tile keeps its Databricks link. The other app draws text in its own fonts, so labels can shift slightly. For PowerPoint, the deck's appendix carries the same editable board.
| Format | Opens in |
|---|---|
draw.io (.drawio) |
draw.io on the web, desktop, Confluence and Jira; imports into Lucidchart |
Visio (.vsdx) |
Visio; imports into Miro, Lucidchart, draw.io and OmniGraffle |
SVG (.svg) |
Figma, Illustrator and Inkscape, as editable vector layers and text |
Excalidraw (.excalidraw) |
Excalidraw, as grouped elements with editable text |
Every download is named for what is in it, so a folder of them stays readable:
databricks-airlines-reference-architecture.pdf, and -platform on the end when
the platform zoom is on. The board carries
(C) Databricks Industry Solutions in its bottom right corner, on screen and in
every export.
The board content translates into sixteen languages: English, French, Spanish,
Chinese, Arabic, Hindi, German, Portuguese, Dutch, Japanese, Italian, Swedish,
Korean, Danish, Finnish and Hebrew. Arabic and Hebrew turn the whole layout
right-to-left. What translates is the content a reader is there to understand,
the zone labels, the captions, the use cases, the team and Genie Agent
descriptions; what does not is the toolbar and menus, which stay in English, and
product and brand names, which are the same word in every language. Each language
is a folder of JSON under translations/, fetched only when it is picked, so the
first load stays small and a language a reader never opens is never downloaded.
The choice carries into the PDF and PowerPoint exports too, so a board read in
Japanese exports a Japanese deck.
The Share control posts the board to LinkedIn, X, Reddit, Facebook, Hacker
News or WhatsApp, with a link straight to the industry and cloud on screen. Every
industry carries its own Open Graph preview image and a small share stub per
cloud, so when one of those links is pasted into Slack, Teams, an email or a
social feed it unfurls with a picture and a title tailored to that industry,
rather than a bare URL. The images and stubs are generated, one per industry and
one per industry-and-cloud, and live under assets/og/ and share/.
Click any box.
| Section | |
|---|---|
| Stage badge | GA, Beta or the preview the box is in |
| What it is | Plain English, no marketing |
| Worth knowing | The thing the name does not tell you |
| Capabilities | What it actually does |
| Learn more | Cloud-specific documentation, the cloud vendor's own page, the product page, and a blog or deep dive |
| Related | The boxes it touches, which highlight on the diagram and are one click away |
A Teams tile reads a little differently, because a team is not a product: the panel names the team, says what it is accountable for, lists its sub-personas and what each one cares about, and then lists the platform surfaces it actually works in with one line each on what it uses them for, alongside the use cases the team is interested in. Those surface and use-case names are clickable, so the panel is a route into the diagram rather than a description beside it.
A use case tile carries its own story: the problem it solves, who benefits (a team from the diagram, clickable), how it is built, the architecture components it uses (each clickable), and links to Databricks customer stories where they exist. On every board the use cases that have a story are shown first.
A source tile is enriched beyond its name: what the system does, who uses it, and the data it produces, split into Batch and Streaming, each stamped with the data shape (structured, semi-structured or unstructured), a typical volume and a cadence. So a reader can see not just that a source feeds the platform but what shape and how much data lands, and how often.
A Genie Agent tile carries the domain it serves as its subtitle, the same way a use case carries its domain, then the governed data sources it reads (clickable chips), the teams it serves, and the top questions it answers in plain language. An AI/BI Dashboard tile names the KPIs and metrics it tracks and the teams that run on it. Both live in their own bands, Genie Agents beside the use cases above the platform, dashboards among the consumers, so the board says who asks the questions and who reads the answers, not just where the data goes.
The three medallion layers open the data model for the industry on screen, so on an airline board that is the airlines model rather than the catalogue root. They also open the launch blog, the Vibe Data Modeling blog and the agent that generates the models.
Every box is reachable by keyboard, and Escape closes the drawer.
The Reference Architecture tab is pinned and cannot be closed. Edit clones the board on screen into a tab of your own: an independent working copy that carries its own industry, cloud, palette, shape, stage and platform zoom, so you can retarget and export it without touching the reference. One clone at a time, so Edit steps aside while a clone is open and returns the moment you close it. Rename a tab by double-clicking its name, and close it with its own x.
What is remembered between visits: theme, palette, platform shape, stage filter, and your tabs, including which one was open. The cloud switch and Platform-only start fresh on every load, because both are how you frame one conversation rather than a preference.
A chat-driven sub-app that lets you describe a customer or use case and receive a new, editable, persisted tab containing a tailored architecture, without touching the reference board. The reference board stays live and unmodified; the generated tab is its own independent copy.
What it does. Type a description in the chat input. The assistant detects the best-fit industry from what you wrote, generates a tab filtered to the components most relevant to that use case, and opens it with edit mode pre-enabled. The tab persists between reloads. The reference board is never switched or mutated by industry detection.
Two-phase, industry-grounded selection. Phase 1 grounds the model on the live generic board catalog and identifies the best-fit industry. Phase 2, only when a known built industry matches, loads that industry's YAML template and re-runs component selection against that industry's own catalog, so industry-specific atoms (for example, banking's core-banking systems or healthcare's EHR vendors) are included in the output. Generic descriptions skip Phase 2.
Industry type-ahead chip. As you type in the chat input, matching industries surface as chips. Picking one switches the Reference board to that industry immediately; it does not generate a tab, it is a separate shortcut to the board's existing industry switch.
Editing a generated architecture. Every generated tab opens with edit mode on (the
✎ Editing: on toggle in the chat header), so it is a working draft you can shape by
hand, not a fixed output. Click any component on the board to edit it in the drawer:
change its title, its caption / description line, the detail text that
explains what it does in this architecture, and its capabilities (comma-separated
tags), or delete it outright. Double-click a tab's name to rename it. Edits are
scoped to that tab and persist between reloads; the toggle is per-tab, so switching to the
Reference board turns editing off and the reference stays read-only. The AI's per-component
usage notes ride along too: hover any component on a generated tab to see why the assistant
included it.
Shared templates, zero extra maintenance. app/ai/index.html sets <base href="../">
so all relative fetches resolve against app/. The assistant reuses the shared
app/architectures/*.yaml templates, app/resources/*.json, app/translations/*, and
supporting JS files from the parent directory. Any new industry template added to
app/architectures/ is automatically available to the AI assistant with no code changes.
Backend. app/ai/app.py is a FastAPI server (not the upstream stdlib main.py). It
serves index.html, exposes /health and POST /generate {system,user,model}->{text},
and mounts the parent app/ subdirectories as static paths so the <base href="../">
fetches resolve. /generate calls a Databricks-hosted Claude Foundation Model serving
endpoint (default databricks-claude-sonnet-5, override via SERVING_ENDPOINT env)
using the injected WorkspaceClient OAuth identity, and no API key is required or stored.
The same 70 industry reference architectures are also exposed as an MCP server, so an agent — Claude Code, Claude Desktop, or any Model Context Protocol client — can list, fetch, and search them programmatically, without the browser app.
It reads app/architectures/*.yaml and app/resources/*.json directly — the same shared
templates the web app uses — so a new industry dropped into architectures/ shows up
through the tools with no code change.
Tools
| Tool | Args | Returns |
|---|---|---|
list_industries |
none | [{id, name, description}] for all 70 industries |
get_architecture |
industry_id |
one industry's full architecture (sources, cloud integrations, pipelines/medallion, consumers, agent use cases) |
search_architectures |
query |
[{id, name, matches}] — industries whose name/description/components mention the query |
list_resources |
kind |
one of the shared maps: accelerators, connectors, links, references |
Run it
cd app/mcp
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python server.py # stdio transportRegister with Claude Code (tools then appear as mcp__arch-explorer__*):
claude mcp add arch-explorer -- "$(pwd)/.venv/bin/python" "$(pwd)/server.py"Or add to a project .mcp.json:
{ "mcpServers": { "arch-explorer": { "command": "/abs/path/app/mcp/.venv/bin/python",
"args": ["/abs/path/app/mcp/server.py"] } } }Use it. Just ask in natural language — the client's own model calls the tools and grounds its answer in the live templates:
- "List the industry reference architectures." →
list_industries - "Show me the banking reference architecture." →
get_architecture("banking") - "Which industries cover fraud detection?" →
search_architectures("fraud") - "What Lakeflow connectors are available?" →
list_resources("connectors") - "Design a real-time fraud architecture for a retail bank" → the model searches the templates, pulls the closest industry, and synthesises a tailored architecture from it.
Retrieval only — it serves the templates as-is. Generating a tailored architecture from a
free-text description is done by the calling model, not a server-side LLM tool. See
app/mcp/README.md for details.
The repository ships an installer that creates the Databricks App and deploys the diagram into it. It needs no catalog, no SQL warehouse and no data access: the app serves one static file, and its service principal reads nothing.
- In your Databricks workspace, choose Workspace -> Create -> Git folder and clone this repository.
- Open
app-installer.ipynbfrom inside that folder. - Run All. The defaults are correct for this path.
- When it finishes it prints a link. The app also appears under Compute -> Apps -> idea.
Run All launches the install as a tagged Databricks job, prints the run URL,
waits for it, then surfaces the app link. The job carries dbx_reference_architecture_agent_installer_*
tags (app, kind, version, status), so its serverless spend is
attributable in system.billing.usage, the same pattern the vibe-modelling
agent installer uses. If the running identity cannot create a job, the notebook
deploys inline instead, so the install still works, just untagged.
A Databricks App has no custom-tag field of its own, so the app's own ongoing
serverless spend is attributed through a serverless usage policy: set widget
4 to a policy id and the installer attaches it on create, and its tags flow to
system.billing.usage.
Widgets, if you want to change something:
| Widget | Default | What it does |
|---|---|---|
01_app_name |
idea |
Name of the Databricks App. Lowercase letters, digits and dashes |
02_source |
Beside this notebook | Where the app files come from. Leave it alone when running from a Git folder |
03_github_repo |
this repo | Only read when 02_source is Download from GitHub |
04_usage_policy_id |
empty | Optional serverless usage policy id to attribute the app's spend. Blank skips it |
Choose Download from GitHub if you imported only the notebook rather than
cloning the repository. It pulls the archive over HTTPS and writes the app/
folder into your workspace home. Two things have to be true for it to work: the
workspace needs outbound internet access, and the repository has to be readable
without a token, which a public repository is. A private repository returns 404
to the anonymous archive request, and some locked-down workspaces have no
outbound internet either, so use the Git folder path in both of those cases.
To upgrade later: Pull on the Git folder, then Run All again. The installer reuses the existing app and redeploys it.
Prerequisites
- Databricks Apps enabled on the workspace (Compute -> Apps)
- Permission to create apps, which workspace admins have by default
Cost. The app runs on its own small serverless compute. Stop it from Compute -> Apps when you are not showing it.
If you would rather not run a notebook:
# 1. put the whole app folder in the workspace (boards, translations,
# resources, assets and vendor all have to come along, not just index.html)
databricks sync app "/Workspace/Users/$USER/idea/app"
# 2. create the app and deploy into it
databricks apps create idea
databricks apps deploy idea --source-code-path "/Workspace/Users/$USER/idea/app"Redeploying after a change is the same two steps without apps create.
Serve the folder with the bundled static server and open the printed URL:
cd app && python3 main.py # http://localhost:8000The boards and translations are fetched at runtime, so serve the folder rather
than opening index.html straight off the filesystem, which browsers block
fetch() on. To hand someone a single file that opens with no server at all, use
the HTML download from the toolbar: it bakes the current board in.
app-installer.ipynb Databricks App installer, Run All
app/
index.html the board: markup, styles, logic, logos; lazy-loads the rest
arch_schema.js the YAML board-descriptor schema, shared by the app and the tools
industry_icons.js the line-art industry glyphs
main.py static server for Databricks Apps, standard library only
app.yaml Databricks App entry point
architectures/ one YAML board per industry (70) plus manifest.json, fetched on demand
resources/ reference tables as JSON: accelerators, connectors, links, references
translations/ one folder per language (English source plus fifteen), fetched on demand
assets/ the Databricks mark, the default share cover, and per-industry cards (assets/og/)
share/ per-industry, per-cloud share stubs (252) that unfurl a preview card
vendor/ js-yaml, the one bundled dependency
ai/ the AI Architecture Assistant sub-app (see app/ai/README.md)
mcp/ MCP server exposing the reference architectures as tools (see app/mcp/README.md)
docs/ the screenshots and the animated export used above
tools/ generators and gates: palettes (palgen.py), product marks (markgen.py),
the installer (build_installer.py), and the height, layout, link, i18n,
export and display checks run before a board ships
A static folder, no build. app/index.html carries the markup, the styles,
the logic and the logos, and loads everything else as data: the seventy
industry boards are YAML descriptors under architectures/, indexed by
architectures/manifest.json and fetched on demand; the reference tables
(accelerators, connectors, links, customer stories) are JSON under resources/;
and each language is a folder of JSON under translations/, fetched only when it
is picked. The one bundled dependency is vendor/js-yaml.min.js, which parses the
board descriptors. There is still no bundler and no package manager, but because
the boards and translations are fetched at runtime the folder is served rather
than opened from a file:// path, which browsers block fetch() on. The HTML
download bakes the current board in, so that single file does open anywhere with
no server.
The model is data. Each board is a plain ARCH object the zones render from,
authored as a readable YAML descriptor and loaded on demand. The reference
material behind each box is a separate table on purpose: the board is what a user
edits, saves and exports, while the product descriptions, stages and links are
fixed facts that have no business being editable. You can export any board as
YAML from the Download menu, edit it offline, and import it back to open your
version in a new tab.
Exports are written by hand. The PDF, PowerPoint and GIF writers are in the
file: object tables and cross-reference offsets for the PDF, the parts and
relationships of an Office package for the PPTX, both carrying the cover, an
index, the board, then a page or slide per item across the Use Cases, Genie
Agents, AI/BI Dashboards and Databricks Apps sections, and a closing slide, with
clickable links throughout. The PPTX then adds an appendix with the board as
editable shapes. Every section is laid out from the same tile data the
on-screen drawer reads, through one deckSections() source of truth, so the deck
and the drawer cannot drift, and each tile is relabelled Branded or Categorized
and picks up the live theme, so a categorized dark board exports a categorized
dark deck. Pulling in a library for each format would be more code than the
formats need, and would put the exports behind a network fetch that a workspace
with no internet egress would fail on. The draw.io, Visio, SVG and Excalidraw
writers and the deck's editable appendix share one walk of the live board,
collectBoard(), so all five carry the same shapes, groups and links.
Colours are solved, not chosen. tools/palgen.py takes a hue recipe per
zone and walks lightness until every foreground clears WCAG AA against the
surface it will actually sit on, in both themes, then emits the CSS.
The animation survives export. A raster export freezes every CSS animation,
so a naive GIF would be twelve identical frames. Each connector's travelling dot
is also expressed as a function of a --dash-t phase variable, which the GIF
encoder walks one step per frame. Only the moving pixels differ between frames,
and the encoder stores the rest as transparent, which is why twelve frames of a
1400px board fit in about 200 KB.
The layout is measured, not tuned. The board is laid out at a fixed design width and then scaled to fit, and the two halves of the platform are balanced after the first measurement and re-measured before the scale is applied. That is why a shape change, a stage filter and a palette switch all land on the same symmetrical geometry instead of drifting a few pixels each time.
Content is written to a measured budget. Because the board is scaled to fit,
content and type size are the same decision: a zone that grows by one tile
shrinks every label on the board. So the pockets are measured rather than
guessed. tools/probe.py injects a probe that clones tiles into a zone until the
board gets taller, which is how the eight-tile Teams pocket is an eight-tile
pocket, and tools/heightgate.py holds every industry to it in all five shapes.
Built by Ashraf Osman & Amr Ali.
The Apache Spark, MLflow, Delta Lake, Unity Catalog, Apache Iceberg and Delta Sharing marks belong to their respective projects and are used to identify those projects.












