Skip to content

About

Interactive explorer for the Databricks Data Intelligence Platform reference architecture — 63 industry-tailored boards, multi-cloud, with an AI assistant that generates tailored architectures from a natural-language description.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Latest commit

 

History

185 Commits

Folders and files

Repository files navigation

Databricks Reference Architecture

An interactive, exportable reference architecture for the Databricks Data Intelligence Platform. No build step and no backend: a static folder that lazy-loads its industry boards, translations and reference data as plain files, translates into sixteen languages, and exports to PDF, PowerPoint, PNG, GIF, a standalone HTML copy, and editable files for draw.io, Lucidchart, Visio, Miro, Figma and Excalidraw. Serve it locally, or deploy it into your own Databricks workspace as an app your teams reach from the workspace navigation.

Open the live board

Databricks Reference Architecture in light theme


What it is

Most reference architectures are a picture. Someone drew it in a diagramming tool, exported a PNG, and pasted it into a deck. It is accurate on the day it is made, it cannot be interrogated, and adapting it to a specific customer means starting again in the source tool, which nobody has.

Databricks Reference Architecture is the same architecture as a live document:

  • Every box is a real product, and clicking it opens what it is, what a customer would not learn from the name, its release stage, the documentation for the cloud you are on, the product page, and related boxes you can jump to.
  • The cloud provider is a switch. Azure, AWS and GCP swap the storage, compute, identity and ingestion services, and every documentation link re-points at that cloud's own docs.
  • The platform is drawn in five shapes, so the same architecture fits a 16:9 slide, a portrait page, or a layout that has to leave room for a third party in the middle.
  • Release stage is a filter. Show GA only for a procurement conversation, or add Beta and the previews for a roadmap one.
  • Industry is a switch too. Sixty-three industries, each one specialising the sources, ingestion, teams, apps, use cases and consumers, with the medallion layers pointing at that industry's own data model. Every use case and every team is a story you can open: click a use case for the problem it solves, who benefits, how it is built, the components it uses and the customer stories that prove it; click a team for its sub-personas and the use cases they care about.
  • The flows are alive. Every connector is a solid arrow with a glowing dot gliding from source to target, each on its own timing, so the board reads as many independent data flows rather than one static picture. The motion is captured in the GIF export too.
  • A guided tour walks a first-time viewer through every control and every zone, and opens the detail panel on a real box so they see what a click gives them. It runs itself once on a first visit and replays any time from the ◎ button in the toolbar.
  • It speaks sixteen languages. A language switch translates the board content into any of sixteen languages, right-to-left for Arabic and Hebrew, while the toolbar stays in English and product and brand names are never translated. Each language is fetched only when it is picked.
  • Every board has a share card. A share button posts the board straight to LinkedIn, X, Reddit and more, and every industry has its own Open Graph preview image, so a pasted link unfurls in Slack, Teams or a social feed with a picture and title tailored to that industry rather than a bare URL.
  • It exports, and the export points back. PDF and PowerPoint open on a cover carrying the title and a link to the exact live board the file was made from, then an index of the sections, the architecture, and one detailed page or slide per item across four sections: Use Cases, Genie Agents, AI/BI Dashboards and Databricks Apps, each carrying the same detail the on-screen drawer shows, and a closing slide that links back to the platform and the live board. The deck follows the theme and the Branded/Categorized choice on screen, so a categorized dark board exports a categorized dark deck. Plus PNG, an animated GIF that keeps the flow moving, and a standalone HTML copy.

Databricks Reference Architecture in dark theme


Why it is useful

1. See the whole of Databricks, and where it fits your architecture

Most people meet Databricks one feature at a time: a notebook here, a job there, Unity Catalog in a governance review. This board puts the entire Data Intelligence Platform in one frame: the sources on the left; ingestion, the medallion layers, unified governance and the agentic layers through the middle; and the consumers, the teams and the cloud services around the outside, with the data flows drawn as live arrows so you can see how one part feeds the next. Every box is a real product, not a label, so clicking one tells you what it is, the thing the name does not give away, what it actually does, and the boxes it connects to. In one screen a newcomer gets the shape of the whole platform, and an architect gets the connections between its parts.

It is not a generic diagram either, and that is the point. The cloud switch redraws the storage, compute, identity and ingestion services as Azure, AWS or GCP and re-points every documentation link at that cloud, so you are looking at the platform on the cloud you actually run. The industry switch respecialises the sources, teams, use cases and consumers for your sector and points the medallion layers at that industry's own data model. And when you need it tied to one specific estate, edit mode, the YAML export and import, and the AI assistant let you retarget the board to a particular customer or use case and export it, as a PDF, a PowerPoint or a standalone page, that shows exactly how Databricks fits into that architecture. It is the difference between a picture of the platform and a picture of your platform.

2. Learn the whole platform, not just present it

Because every box carries its own facts and links, the board doubles as a way to learn Databricks rather than only to show it. Open any product and the Learn more section takes you to the cloud-specific documentation for the cloud you are on (the Azure Databricks docs on Azure, docs.databricks.com on AWS and GCP), the product page, and a blog or deep dive, so each box is a launch point into the official material rather than a dead end. The release-stage filter teaches what is generally available today against what is still in preview or beta. The use cases spell out the problem each one solves, who benefits, how it is built and which components it touches, with links to real customer stories, so you see how the features are used and not only what they are named. The Genie Agents and AI/BI dashboards show the question-and-answer layer that sits on top. There is a Databricks YouTube channel section with curated videos verified against the official channel for when you would rather watch than read, and the whole board translates into sixteen languages, so a team can learn in its own. A new joiner, a partner or a customer can explore the platform at their own pace, from the first box to the deepest link, without a slide deck in between.


What it shows

Sources on the left, the platform through the middle, consumers on the right, with the teams who use it and the cloud it runs on wrapped around the outside.

Zone What sits there
Sources Structured, semi-structured, unstructured, streaming and IoT, external and partner data, and federation sources reached without copying
Cloud and 3rd-party ingestion The cloud's own ETL services, and third-party ELT and streaming brokers that land data alongside the platform's native ingestion
Platform Ingest, Agentic Apps, Agentic Work, Unified Governance, Agentic Data, Open Infrastructure and the medallion layers, drawn as one outline traced by a moving ring
Teams The business and technical teams the platform is built for. A tile is a team, never a job title, and the roles inside it and the surfaces they touch are in the side panel
Consumers BI and productivity tools, MCP and APIs, published data products, partners and platforms, operational systems, four AI/BI Dashboards and four Databricks Apps, and the agent harnesses that arrive from outside
Use cases and Genie Agents What the platform is used for, in the band above the platform: ten use cases and four Genie Agents on every board, each Genie Agent labelled with the domain it serves, industry-neutral by default
Cloud services and integrations The account's own storage, compute, key vault, catalog, identity and observability services

How to read it

Signal Meaning
Solid arrow Data moves along it, and a glowing dot glides from source to target in the direction it actually flows. Each arrow runs on its own timing, so the board reads as many independent flows rather than one synchronised pulse
Dashed zone outline A grouping, not a boundary that data crosses
The ring around the platform One continuous outline: the platform is one product, not a stack of separate ones
Colour Identifies the zone, never the status. Every zone keeps its own hue in every palette and both themes

The controls

Fourteen controls, left to right, sitting in the header above the diagram, plus platform zoom on the diagram itself.

Control What it does
Share feedback Opens a feedback form in a new tab, for a comment on the reference architecture itself
Share Posts the board on screen to LinkedIn, X, Reddit and more; the link points straight to the industry showing, and unfurls with that industry's own preview card
Industry Sixty-three industries plus Standard Reference Architecture, searchable, every entry on one line. Specialises everything outside the platform: sources, ingestion, teams, apps, use cases and consumers. The platform itself and the cloud services band do not change
Display: Branded / Categorized Shows every product by its commercial brand name (Branded), or rooted back to its generic category (Categorized, the default) for a room that knows the category but not the product, so SABRE reads as Airline Reservation System. The choice applies everywhere at once, the tiles, the detail drawers and every download, and it is remembered between visits
Cloud Azure, AWS, GCP. Swaps the cloud services band, the cloud ETL tiles and the federation sources, and re-points every documentation link at that cloud's own docs, including the Microsoft Learn pages on Azure
Dark / Light Follows the operating system by default, and remembers an explicit choice. Downloads follow whatever is on screen
Palette Thirteen colour schemes in three groups
Style Five platform shapes
Stage Filters the platform box by release stage
Download PDF, PowerPoint with an editable appendix, PNG, GIF, HTML, and editable draw.io, Visio, SVG and Excalidraw files. Its Architecture descriptor section exports the board on screen, the reference or any industry, as an editable YAML file, and imports one back to open your own version in a new tab
Details panel dock Pins the detail drawer to the side so it stays open while you click from box to box, instead of overlaying the board each time
Language Translates the board content into any of sixteen languages, right-to-left for Arabic and Hebrew. The toolbar and menus stay in English, and product and brand names are never translated
Tour (the ◎ at the end of the toolbar) A guided walk-through that spotlights each control and each zone in turn, opens the detail panel on a real box so you see what a click gives you, and runs itself once on a first visit. Replayable any time
Zoom (on the platform heading, not the toolbar) Zooms into the platform on its own: hides sources, consumers, the apps band and the cloud services. Click again to restore. Exports respect it, and the button itself never appears in one

Everything that acts on a diagram goes inert on a tab that has no diagram yet, so the toolbar cannot be used against nothing. Dark/Light and Cloud stay live, because they are global and are the two things worth setting before a diagram exists.

Industry, cloud, shape, palette, theme and platform zoom are each a URL parameter, so a board can be linked to in the state it was read in: ?industry=airlines&cloud=aws&shape=h90&pal=nordic&theme=light&platform=1. That is the link the PDF and PowerPoint covers carry. The Branded/Categorized choice is remembered locally rather than in the link, so a shared URL opens in the reader's own preferred naming.

Industry: seventy models, and the same board size for each

The industry list follows the Databricks Industry Data Models catalogue and extends it: airlines is the reference, and sixty-two more are authored to the same depth, spanning finance, healthcare and life sciences, public sector, manufacturing and energy, retail and consumer, media, technology, professional services and more. Picking one rewrites the four zones outside the platform, and the medallion layers start pointing at that industry's own folder in the repository.

Every industry is written to the same schema as the airlines reference: real, current vendor systems in the sources and ingestion, teams broken into sub-personas, four apps and ten use cases, and each use case carrying a problem statement, its beneficiary, how it is built, the architecture components it touches, and links to matching Databricks customer stories. Use cases with a story sort ahead of those without, so the board leads with proof.

An industry board is held to the reference board's height, which matters more than it sounds: the board is scaled to fit, so one industry carrying a few more tiles than the rest would render every label in that industry smaller. Teams and ingestion each get a fixed pocket, eight tiles and three groups, and industry content is written to that budget rather than allowed to grow past it. tools/heightgate.py measures every industry against the reference in all five shapes and fails on a board that is taller, a label that is clipped, or a label that only fits by wrapping.

Every industry, one click away

Each name opens the live board specialised for that industry: its own sources and ingestion, its own teams, its own four apps and ten use cases, its own consumers, and the medallion layers pointing at that industry's data model. Add &cloud=aws or &cloud=gcp to any link to open it on that cloud; the default is Azure.

Sector Industries
Financial services Banking · Capital Markets (Buy-Side) · Capital Markets (Sell-Side) · Payments & Fintech · Wealth Management · Mortgage & Lending · Exchanges · Market Data & Analytics Providers · Crypto & Digital Assets
Insurance Insurance (P&C) · Life Insurance · Health Insurance
Healthcare & life sciences Healthcare · Digital Health · Pharmaceuticals · Pharmacy & PBM · Clinical Trials · Genomics & Biotech · Diagnostics & Labs · Medical Devices
Public sector & education Public Sector · Public Safety · Education · EdTech · NGO & Non-Profit
Manufacturing & industrial Manufacturing · Automotive · Aerospace & Space · Semiconductors · Chemical Manufacturing · Paper & Packaging · Construction · Computer & Electronics · Industrial Machinery
Energy & resources Electric Power Generation · Electric Utility (T&D) · Oil & Gas (Upstream) · Oil & Gas (Midstream) · Oil & Gas (Downstream & Refining) · Renewables & Cleantech · Mining · Water Utilities · Waste Management
Retail & consumer Retail · E-Commerce · Grocery · Consumer Goods · Apparel & Fashion · Food & Beverage · Restaurants · Wholesale & Distribution
Travel, transport & logistics Airlines · Travel & Hospitality · Transport & Logistics · Shipping & Ports · Rail & Transit
Technology, media & telecom Software & Technology · Cybersecurity · Data Centers & Cloud · Telecommunications · Media & Broadcasting · Advertising · Gaming · Sports & Entertainment
Professional & business services Professional Services · Legal · Staffing & HR · Real Estate
Agriculture Agriculture · AgTech

Sixty-three industries in all, plus the industry-neutral Standard Reference Architecture that every board starts from.

Style: five platform shapes

The five platform shapes

Shape Why it exists
Z The default. Ingest reaches left over the sources, serving reaches right over the consumers
S The Z mirrored, for when the story runs right to left
T Both arms on top, for a wide slide with a short middle
T180 Both arms underneath
H90 A full I-beam: three rows with split arms and pockets inside the notches, which is what leaves room for the cloud and third-party ingestion to sit inside the shape rather than beside it

Every shape is measured, not eyeballed: the two pockets are equalised to the taller one, so the arms stay symmetrical to the pixel in all five.

H90 works differently enough from the other four to be worth spelling out. Each pocket splits into two labelled boxes side by side, Cloud ETL beside 3rd Party and Business beside Technical, and each box runs its tiles down a single column. Both pockets sit inside their own arm column rather than spanning the middle, so the crossbar of the I stays clear, and the governance band renders there instead of in the lower block: the narrow middle of the shape carries the platform's control plane rather than being a spacer. The arms are wider here than in the other shapes to fit those boxes, and the design canvas is wider to match, which is free because the fit in this shape is bound by height rather than width.

H90, the full board

Palette: thirteen schemes, solved rather than picked

Six of the thirteen palettes

Group Palettes
Neutral Spectrum (default), Mono (print safe), Muted (low chroma), Nordic (cool calm)
Coloured Ocean (analogous), Earth (warm neutral), Sunset (warm shift), Berry (cool warm)
Loud Solid (filled), Jewel (deep), Vivid (projector), Pop (playful), Neon (maximum)

tools/palgen.py generates all of them. For every zone hue it walks lightness until four values clear WCAG AA against the surface each one actually sits on, in both themes: the zone fill, the chip tile inside it, the border, and the ink on top. Nothing here is hand-picked, which is why Neon is legible and Mono survives a black and white printer.

Spectrum keeps the chips white so the reference reads as a document. Every other palette tints the chips too, so a palette choice is visible in every box rather than in the labels alone.

Stage: filter by release stage

Stage
GA Available now
Public Preview
Beta
Private Preview
Coming soon

Each row carries the number of boxes it would show, and switching a stage off removes those boxes from the platform, closes the gap and re-traces the ring. Unstaged boxes always stay. An st value the app does not recognise is treated as unstaged rather than quietly filed as available, because the conservative reading is the honest one in front of a customer.

Download

Format What you get
PDF A cover with the title, the industry and cloud, the sentence the board leads with, and clickable links to the exact live board and to the industry's data model; an index page listing the sections; the architecture as a pixel-exact image of the board, in the current theme and palette; then four section breaks, each with one page per item, Use Cases (problem, beneficiary, build, components, story links), Genie Agents (the domain it serves, the data it reads, the teams and its top questions), AI/BI Dashboards (the KPIs it tracks and the teams that run on it) and Databricks Apps; then a closing page linking to the platform and the live board. All in the theme on screen and in Branded or Categorized names to match the board
PowerPoint The same pages, as native slides: a cover, an index slide, a board slide carrying a pixel-exact image of the board, then the Use Cases, Genie Agents, AI/BI Dashboards and Databricks Apps sections a slide per item, and a closing slide. An Editable Architecture appendix follows: a break slide, then the board again as grouped native shapes you can ungroup and edit in PowerPoint, Keynote or Google Slides. The deck follows the theme and the Branded/Categorized choice on screen, so a categorized dark board exports a categorized dark deck
PNG 2x raster of the current view
GIF A looping animation, 1400px wide, twelve frames, that keeps the travelling dashes and the platform ring moving. Roughly 200 KB, because only the moving pixels are stored per frame, in the palette and theme on screen
HTML A standalone copy of the page with your current choices baked in, which opens anywhere with no server

Edit in another app. These downloads rebuild the current view as grouped, editable shapes instead of a picture. Every zone, group box and tile is a group you can ungroup; labels are real text, arrows are connectors, product logos are crisp images and each tile keeps its Databricks link. The other app draws text in its own fonts, so labels can shift slightly. For PowerPoint, the deck's appendix carries the same editable board.

Format Opens in
draw.io (.drawio) draw.io on the web, desktop, Confluence and Jira; imports into Lucidchart
Visio (.vsdx) Visio; imports into Miro, Lucidchart, draw.io and OmniGraffle
SVG (.svg) Figma, Illustrator and Inkscape, as editable vector layers and text
Excalidraw (.excalidraw) Excalidraw, as grouped elements with editable text

Every download is named for what is in it, so a folder of them stays readable: databricks-airlines-reference-architecture.pdf, and -platform on the end when the platform zoom is on. The board carries (C) Databricks Industry Solutions in its bottom right corner, on screen and in every export.

The animated GIF export

Language: sixteen translations

The board content translates into sixteen languages: English, French, Spanish, Chinese, Arabic, Hindi, German, Portuguese, Dutch, Japanese, Italian, Swedish, Korean, Danish, Finnish and Hebrew. Arabic and Hebrew turn the whole layout right-to-left. What translates is the content a reader is there to understand, the zone labels, the captions, the use cases, the team and Genie Agent descriptions; what does not is the toolbar and menus, which stay in English, and product and brand names, which are the same word in every language. Each language is a folder of JSON under translations/, fetched only when it is picked, so the first load stays small and a language a reader never opens is never downloaded. The choice carries into the PDF and PowerPoint exports too, so a board read in Japanese exports a Japanese deck.

The language menu: sixteen languages, each in its own script

Share: a preview card per industry

The Share control posts the board to LinkedIn, X, Reddit, Facebook, Hacker News or WhatsApp, with a link straight to the industry and cloud on screen. Every industry carries its own Open Graph preview image and a small share stub per cloud, so when one of those links is pasted into Slack, Teams, an email or a social feed it unfurls with a picture and a title tailored to that industry, rather than a bare URL. The images and stubs are generated, one per industry and one per industry-and-cloud, and live under assets/og/ and share/.

The share menu: a link straight to the industry and cloud on screen


The detail drawer

Click any box.

The detail drawer, opened on Unity Catalog

Section
Stage badge GA, Beta or the preview the box is in
What it is Plain English, no marketing
Worth knowing The thing the name does not tell you
Capabilities What it actually does
Learn more Cloud-specific documentation, the cloud vendor's own page, the product page, and a blog or deep dive
Related The boxes it touches, which highlight on the diagram and are one click away

A Teams tile reads a little differently, because a team is not a product: the panel names the team, says what it is accountable for, lists its sub-personas and what each one cares about, and then lists the platform surfaces it actually works in with one line each on what it uses them for, alongside the use cases the team is interested in. Those surface and use-case names are clickable, so the panel is a route into the diagram rather than a description beside it.

A use case tile carries its own story: the problem it solves, who benefits (a team from the diagram, clickable), how it is built, the architecture components it uses (each clickable), and links to Databricks customer stories where they exist. On every board the use cases that have a story are shown first.

A source tile is enriched beyond its name: what the system does, who uses it, and the data it produces, split into Batch and Streaming, each stamped with the data shape (structured, semi-structured or unstructured), a typical volume and a cadence. So a reader can see not just that a source feeds the platform but what shape and how much data lands, and how often.

An enriched source drawer, opened on Business Applications

A Genie Agent tile carries the domain it serves as its subtitle, the same way a use case carries its domain, then the governed data sources it reads (clickable chips), the teams it serves, and the top questions it answers in plain language. An AI/BI Dashboard tile names the KPIs and metrics it tracks and the teams that run on it. Both live in their own bands, Genie Agents beside the use cases above the platform, dashboards among the consumers, so the board says who asks the questions and who reads the answers, not just where the data goes.

A Genie Agent drawer, opened on the Customer & Revenue Agent

An AI/BI Dashboard drawer, opened on Revenue & Growth

The three medallion layers open the data model for the industry on screen, so on an airline board that is the airlines model rather than the catalogue root. They also open the launch blog, the Vibe Data Modeling blog and the agent that generates the models.

Every box is reachable by keyboard, and Escape closes the drawer.


Tabs

The Reference Architecture tab is pinned and cannot be closed. Edit clones the board on screen into a tab of your own: an independent working copy that carries its own industry, cloud, palette, shape, stage and platform zoom, so you can retarget and export it without touching the reference. One clone at a time, so Edit steps aside while a clone is open and returns the moment you close it. Rename a tab by double-clicking its name, and close it with its own x.

What is remembered between visits: theme, palette, platform shape, stage filter, and your tabs, including which one was open. The cloud switch and Platform-only start fresh on every load, because both are how you frame one conversation rather than a preference.


AI Architecture Assistant (app/ai/)

A chat-driven sub-app that lets you describe a customer or use case and receive a new, editable, persisted tab containing a tailored architecture, without touching the reference board. The reference board stays live and unmodified; the generated tab is its own independent copy.

The AI Architecture Assistant, docked over the reference board

What it does. Type a description in the chat input. The assistant detects the best-fit industry from what you wrote, generates a tab filtered to the components most relevant to that use case, and opens it with edit mode pre-enabled. The tab persists between reloads. The reference board is never switched or mutated by industry detection.

Two-phase, industry-grounded selection. Phase 1 grounds the model on the live generic board catalog and identifies the best-fit industry. Phase 2, only when a known built industry matches, loads that industry's YAML template and re-runs component selection against that industry's own catalog, so industry-specific atoms (for example, banking's core-banking systems or healthcare's EHR vendors) are included in the output. Generic descriptions skip Phase 2.

Industry type-ahead chip. As you type in the chat input, matching industries surface as chips. Picking one switches the Reference board to that industry immediately; it does not generate a tab, it is a separate shortcut to the board's existing industry switch.

Editing a generated architecture. Every generated tab opens with edit mode on (the ✎ Editing: on toggle in the chat header), so it is a working draft you can shape by hand, not a fixed output. Click any component on the board to edit it in the drawer: change its title, its caption / description line, the detail text that explains what it does in this architecture, and its capabilities (comma-separated tags), or delete it outright. Double-click a tab's name to rename it. Edits are scoped to that tab and persist between reloads; the toggle is per-tab, so switching to the Reference board turns editing off and the reference stays read-only. The AI's per-component usage notes ride along too: hover any component on a generated tab to see why the assistant included it.

Shared templates, zero extra maintenance. app/ai/index.html sets <base href="../"> so all relative fetches resolve against app/. The assistant reuses the shared app/architectures/*.yaml templates, app/resources/*.json, app/translations/*, and supporting JS files from the parent directory. Any new industry template added to app/architectures/ is automatically available to the AI assistant with no code changes.

Backend. app/ai/app.py is a FastAPI server (not the upstream stdlib main.py). It serves index.html, exposes /health and POST /generate {system,user,model}->{text}, and mounts the parent app/ subdirectories as static paths so the <base href="../"> fetches resolve. /generate calls a Databricks-hosted Claude Foundation Model serving endpoint (default databricks-claude-sonnet-5, override via SERVING_ENDPOINT env) using the injected WorkspaceClient OAuth identity, and no API key is required or stored.


MCP server (app/mcp/)

The same 70 industry reference architectures are also exposed as an MCP server, so an agent — Claude Code, Claude Desktop, or any Model Context Protocol client — can list, fetch, and search them programmatically, without the browser app.

It reads app/architectures/*.yaml and app/resources/*.json directly — the same shared templates the web app uses — so a new industry dropped into architectures/ shows up through the tools with no code change.

Tools

Tool Args Returns
list_industries none [{id, name, description}] for all 70 industries
get_architecture industry_id one industry's full architecture (sources, cloud integrations, pipelines/medallion, consumers, agent use cases)
search_architectures query [{id, name, matches}] — industries whose name/description/components mention the query
list_resources kind one of the shared maps: accelerators, connectors, links, references

Run it

cd app/mcp
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
.venv/bin/python server.py            # stdio transport

Register with Claude Code (tools then appear as mcp__arch-explorer__*):

claude mcp add arch-explorer -- "$(pwd)/.venv/bin/python" "$(pwd)/server.py"

Or add to a project .mcp.json:

{ "mcpServers": { "arch-explorer": { "command": "/abs/path/app/mcp/.venv/bin/python",
                                     "args": ["/abs/path/app/mcp/server.py"] } } }

Use it. Just ask in natural language — the client's own model calls the tools and grounds its answer in the live templates:

  • "List the industry reference architectures." → list_industries
  • "Show me the banking reference architecture." → get_architecture("banking")
  • "Which industries cover fraud detection?" → search_architectures("fraud")
  • "What Lakeflow connectors are available?" → list_resources("connectors")
  • "Design a real-time fraud architecture for a retail bank" → the model searches the templates, pulls the closest industry, and synthesises a tailored architecture from it.

Retrieval only — it serves the templates as-is. Generating a tailored architecture from a free-text description is done by the calling model, not a server-side LLM tool. See app/mcp/README.md for details.


How to deploy

Option 1: the installer notebook (recommended)

The repository ships an installer that creates the Databricks App and deploys the diagram into it. It needs no catalog, no SQL warehouse and no data access: the app serves one static file, and its service principal reads nothing.

  1. In your Databricks workspace, choose Workspace -> Create -> Git folder and clone this repository.
  2. Open app-installer.ipynb from inside that folder.
  3. Run All. The defaults are correct for this path.
  4. When it finishes it prints a link. The app also appears under Compute -> Apps -> idea.

Run All launches the install as a tagged Databricks job, prints the run URL, waits for it, then surfaces the app link. The job carries dbx_reference_architecture_agent_installer_* tags (app, kind, version, status), so its serverless spend is attributable in system.billing.usage, the same pattern the vibe-modelling agent installer uses. If the running identity cannot create a job, the notebook deploys inline instead, so the install still works, just untagged.

A Databricks App has no custom-tag field of its own, so the app's own ongoing serverless spend is attributed through a serverless usage policy: set widget 4 to a policy id and the installer attaches it on create, and its tags flow to system.billing.usage.

Widgets, if you want to change something:

Widget Default What it does
01_app_name idea Name of the Databricks App. Lowercase letters, digits and dashes
02_source Beside this notebook Where the app files come from. Leave it alone when running from a Git folder
03_github_repo this repo Only read when 02_source is Download from GitHub
04_usage_policy_id empty Optional serverless usage policy id to attribute the app's spend. Blank skips it

Choose Download from GitHub if you imported only the notebook rather than cloning the repository. It pulls the archive over HTTPS and writes the app/ folder into your workspace home. Two things have to be true for it to work: the workspace needs outbound internet access, and the repository has to be readable without a token, which a public repository is. A private repository returns 404 to the anonymous archive request, and some locked-down workspaces have no outbound internet either, so use the Git folder path in both of those cases.

To upgrade later: Pull on the Git folder, then Run All again. The installer reuses the existing app and redeploys it.

Prerequisites

  • Databricks Apps enabled on the workspace (Compute -> Apps)
  • Permission to create apps, which workspace admins have by default

Cost. The app runs on its own small serverless compute. Stop it from Compute -> Apps when you are not showing it.

Option 2: the CLI

If you would rather not run a notebook:

# 1. put the whole app folder in the workspace (boards, translations,
#    resources, assets and vendor all have to come along, not just index.html)
databricks sync app "/Workspace/Users/$USER/idea/app"

# 2. create the app and deploy into it
databricks apps create idea
databricks apps deploy idea --source-code-path "/Workspace/Users/$USER/idea/app"

Redeploying after a change is the same two steps without apps create.

Option 3: no Databricks at all

Serve the folder with the bundled static server and open the printed URL:

cd app && python3 main.py     # http://localhost:8000

The boards and translations are fetched at runtime, so serve the folder rather than opening index.html straight off the filesystem, which browsers block fetch() on. To hand someone a single file that opens with no server at all, use the HTML download from the toolbar: it bakes the current board in.


Repository layout

app-installer.ipynb        Databricks App installer, Run All
app/
  index.html               the board: markup, styles, logic, logos; lazy-loads the rest
  arch_schema.js           the YAML board-descriptor schema, shared by the app and the tools
  industry_icons.js        the line-art industry glyphs
  main.py                  static server for Databricks Apps, standard library only
  app.yaml                 Databricks App entry point
  architectures/           one YAML board per industry (70) plus manifest.json, fetched on demand
  resources/               reference tables as JSON: accelerators, connectors, links, references
  translations/            one folder per language (English source plus fifteen), fetched on demand
  assets/                  the Databricks mark, the default share cover, and per-industry cards (assets/og/)
  share/                   per-industry, per-cloud share stubs (252) that unfurl a preview card
  vendor/                  js-yaml, the one bundled dependency
  ai/                      the AI Architecture Assistant sub-app (see app/ai/README.md)
  mcp/                     MCP server exposing the reference architectures as tools (see app/mcp/README.md)
docs/                      the screenshots and the animated export used above
tools/                     generators and gates: palettes (palgen.py), product marks (markgen.py),
                           the installer (build_installer.py), and the height, layout, link, i18n,
                           export and display checks run before a board ships

How it is built

A static folder, no build. app/index.html carries the markup, the styles, the logic and the logos, and loads everything else as data: the seventy industry boards are YAML descriptors under architectures/, indexed by architectures/manifest.json and fetched on demand; the reference tables (accelerators, connectors, links, customer stories) are JSON under resources/; and each language is a folder of JSON under translations/, fetched only when it is picked. The one bundled dependency is vendor/js-yaml.min.js, which parses the board descriptors. There is still no bundler and no package manager, but because the boards and translations are fetched at runtime the folder is served rather than opened from a file:// path, which browsers block fetch() on. The HTML download bakes the current board in, so that single file does open anywhere with no server.

The model is data. Each board is a plain ARCH object the zones render from, authored as a readable YAML descriptor and loaded on demand. The reference material behind each box is a separate table on purpose: the board is what a user edits, saves and exports, while the product descriptions, stages and links are fixed facts that have no business being editable. You can export any board as YAML from the Download menu, edit it offline, and import it back to open your version in a new tab.

Exports are written by hand. The PDF, PowerPoint and GIF writers are in the file: object tables and cross-reference offsets for the PDF, the parts and relationships of an Office package for the PPTX, both carrying the cover, an index, the board, then a page or slide per item across the Use Cases, Genie Agents, AI/BI Dashboards and Databricks Apps sections, and a closing slide, with clickable links throughout. The PPTX then adds an appendix with the board as editable shapes. Every section is laid out from the same tile data the on-screen drawer reads, through one deckSections() source of truth, so the deck and the drawer cannot drift, and each tile is relabelled Branded or Categorized and picks up the live theme, so a categorized dark board exports a categorized dark deck. Pulling in a library for each format would be more code than the formats need, and would put the exports behind a network fetch that a workspace with no internet egress would fail on. The draw.io, Visio, SVG and Excalidraw writers and the deck's editable appendix share one walk of the live board, collectBoard(), so all five carry the same shapes, groups and links.

Colours are solved, not chosen. tools/palgen.py takes a hue recipe per zone and walks lightness until every foreground clears WCAG AA against the surface it will actually sit on, in both themes, then emits the CSS.

The animation survives export. A raster export freezes every CSS animation, so a naive GIF would be twelve identical frames. Each connector's travelling dot is also expressed as a function of a --dash-t phase variable, which the GIF encoder walks one step per frame. Only the moving pixels differ between frames, and the encoder stores the rest as transparent, which is why twelve frames of a 1400px board fit in about 200 KB.

The layout is measured, not tuned. The board is laid out at a fixed design width and then scaled to fit, and the two halves of the platform are balanced after the first measurement and re-measured before the scale is applied. That is why a shape change, a stage filter and a palette switch all land on the same symmetrical geometry instead of drifting a few pixels each time.

Content is written to a measured budget. Because the board is scaled to fit, content and type size are the same decision: a zone that grows by one tile shrinks every label on the board. So the pockets are measured rather than guessed. tools/probe.py injects a probe that clones tiles into a zone until the board gets taller, which is how the eight-tile Teams pocket is an eight-tile pocket, and tools/heightgate.py holds every industry to it in all five shapes.


Credits

Built by Ashraf Osman & Amr Ali.

The Apache Spark, MLflow, Delta Lake, Unity Catalog, Apache Iceberg and Delta Sharing marks belong to their respective projects and are used to identify those projects.

About

Interactive explorer for the Databricks Data Intelligence Platform reference architecture — 63 industry-tailored boards, multi-cloud, with an AI assistant that generates tailored architectures from a natural-language description.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages