VoltarBack
Proposta Técnica · On-Premise · Stack Open SourceTechnical Proposal · On-Premise · Open Source Stack Proposta Técnica · Microsoft Azure · Stack Open SourceTechnical Proposal · Microsoft Azure · Open Source Stack Proposta Técnica · Oracle OCI · Serviços GerenciadosTechnical Proposal · Oracle OCI · Managed Services

Data Lake
On-Premise
On-Premise
Data Lake
Data Lake
na Azure
Data Lake
on Azure
Data Lake
na OCI
Data Lake
on OCI

Arquitetura moderna de dados on-premise baseada em ferramentas 100% open source: do banco Oracle transacional até dashboards interativos em tempo real, passando por Airflow, Spark, MinIO e Apache Pinot.Modern on-premise data architecture built on 100% open source tooling: from Oracle transactional data to real-time interactive dashboards, via Airflow, Spark, MinIO, and Apache Pinot. Exatamente o mesmo stack open source, agora sobre a Azure: Airflow, Spark, MinIO, Pinot e Superset rodando em AKS. Zero licenças, zero reescrita e nenhum lock-in — o dia em que quiser sair, o mesmo Helm chart sobe em qualquer outra nuvem ou de volta no seu datacenter.The exact same open source stack, now on Azure: Airflow, Spark, MinIO, Pinot, and Superset running on AKS. Zero licenses, zero rewrite, and no lock-in — the day you want out, the same Helm chart runs on any other cloud or back in your own datacenter. Aqui a premissa se inverte: em vez de carregar o stack open source para a nuvem, aproveita-se o que a OCI já entrega pronto. O Oracle deixa de ser uma fonte externa e passa a viver ao lado do lake, e boa parte da operação sai das suas mãos — em troca de uma dependência bem maior de um fornecedor só.Here the premise inverts: instead of carrying the open source stack to the cloud, you lean on what OCI already provides. Oracle stops being an external source and comes to live next to the lake, and much of the operational burden leaves your hands — in exchange for a far deeper dependency on a single vendor.

On-Premise Oracle → MinIO → Pinot Apache Airflow Apache Spark MinIO S3 Apache Pinot Apache Superset R$ 0 em licenças Oracle OCI Autonomous DB / Exadata Object Storage (S3 API) OCI Data Integration Autonomous Data Warehouse Oracle Analytics Cloud Licenças: sob contratoLicenses: under contract Microsoft Azure AKS Oracle → MinIO → Pinot Apache Airflow Apache Spark MinIO on Managed Disks Apache Pinot Apache Superset R$ 0 em licençasR$ 0 in licenses
3
servidores físicosphysical servers
3
node pools no AKSAKS node pools
1
node pool no OKEOKE node pool
6+
ferramentas open sourceopen source tools
6+
ferramentas open sourceopen source tools
2
componentes ainda open sourcecomponents still open source
<100ms
latência de query PinotPinot query latency
13w
roadmap de implantaçãodeployment roadmap
8w
roadmap de implantaçãodeployment roadmap
6w
roadmap de implantaçãodeployment roadmap
01

Visão da
arquitetura
Architecture
overview

A arquitetura extrai dados do Oracle e os entrega até o analista consultando dados em milissegundos no Superset.The architecture extracts data from Oracle and delivers it all the way to analysts querying in milliseconds on Superset. A mesma arquitetura, agora sobre a Azure. O caminho do dado é idêntico — o que muda é onde os contêineres rodam.The same architecture, now on Azure. The data path is identical — what changes is where the containers run.

Diagrama: End-to-End · On-PremiseDiagram: End-to-End · On-Premise Diagrama: End-to-End · Azure (AKS)Diagram: End-to-End · Azure (AKS) Diagrama: End-to-End · Oracle OCIDiagram: End-to-End · Oracle OCI
FONTE ORQUESTRAÇÃO ARMAZENAMENTO ANALYTICS CONSUMO 🗄️ Oracle DB Transacional · JDBC SQL ⚙️ Airflow Orquestração de DAGs 🔥 Spark Transformações 🥉 Bronze Raw · JSON 🥈 Silver Limpo · Parquet 🥇 Gold Agregado · Parquet MinIO: Data Lake MinIO on AKS · Managed Disks Segment Load Apache Pinot Controller · Broker Server · Segments OLAP <100ms Deep → MinIO SQL · REST API SQL 📊 Superset Dashboards · BI Alertas · Reports 👥 Analistas FONTE INGESTÃO ARMAZENAMENTO ANALYTICS + BI OCI TENANCY · VCN PRIVADA 🗄️ Autonomous DB ou Exadata Cloud CENÁRIO A · migrado mesma VCN · sem WAN latência baixa 🏢 Oracle on-prem CENÁRIO B permanece no seu datacenter FastConnect atravessa WAN ⚙️ Data Integration ETL gerenciado (substitui Airflow + Spark) OCI Object Storage API S3 · sem MinIO 🥉 Bronze Raw 🥈 Silver Parquet 🥇 Gold Parquet load Autonomous DW OLAP gerenciado substitui Pinot auto-tuning · auto-scale 📊 Analytics Cloud Dashboards · BI substitui Superset 👥 Analistas verde = ainda portável (API S3) vermelho = proprietário · reescrita para sair
Premissa de design: Design premise: Todo o stack é on-premise: três servidores físicos dentro da sua infraestrutura. Nenhuma dependência de cloud pública. Pinot usa MinIO como deep storage, mantendo os segmentos no mesmo datacenter. The entire stack is on-premise: three physical servers within your own infrastructure. No public cloud dependency. Pinot uses MinIO as deep storage, keeping all segments in the same datacenter.
Premissa de design: Design premise: O diagrama acima não muda. É exatamente o mesmo stack — só o chão embaixo dele é diferente: em vez de três servidores físicos, três node pools num cluster AKS. Nenhum serviço PaaS proprietário entra no caminho dos dados, então não há reescrita de DAG, de job Spark ou de dashboard. Pinot continua usando MinIO como deep storage, agora sobre Azure Managed Disks. The diagram above does not change. It is the exact same stack — only the floor beneath it differs: instead of three physical servers, three node pools in an AKS cluster. No proprietary PaaS service sits in the data path, so there is no DAG, Spark job, or dashboard to rewrite. Pinot still uses MinIO as deep storage, now backed by Azure Managed Disks.
Premissa de design: Design premise: Aqui o diagrama muda de verdade — é a única das três opções em que isso acontece. Airflow, Spark, Pinot e Superset saem e dão lugar a serviços gerenciados da Oracle. O que se ganha: quase nada para operar, e o Oracle deixa de ser fonte externa. O que se perde: a portabilidade que sustenta as outras duas abas. Só o Object Storage continua neutro, porque fala API S3 — é o único componente que você conseguiria levar embora sem reescrever. Here the diagram genuinely changes — the only one of the three options where it does. Airflow, Spark, Pinot, and Superset step aside for Oracle managed services. What you gain: almost nothing left to operate, and Oracle stops being an external source. What you lose: the portability that holds up the other two tabs. Only Object Storage stays neutral, because it speaks the S3 API — it is the one component you could take with you without a rewrite.
02

A stack The stack

Oracle como fonte, Airflow orquestrando cada passo, Spark transformando em escala, MinIO armazenando em três camadas, Pinot respondendo em menos de 100ms, Superset entregando ao analista. Do dado bruto ao dashboard, sem dependência de cloud pública.Oracle as source, Airflow orchestrating each step, Spark transforming at scale, MinIO storing across three layers, Pinot responding in under 100ms, Superset delivering to the analyst. From raw data to dashboard, with no public cloud dependency. Os mesmos seis componentes, as mesmas versões, os mesmos arquivos de configuração — empacotados em Helm charts e rodando em AKS. A Azure entra como infraestrutura (computação, disco, rede, identidade), nunca como parte do pipeline.The same six components, same versions, same config files — packaged as Helm charts and running on AKS. Azure comes in as infrastructure (compute, disk, network, identity), never as part of the pipeline itself. Dos seis componentes originais, quatro saem. Ficam o Object Storage — que fala API S3 e por isso continua neutro — e o modelo Medallion, que é desenho de dados, não ferramenta. O resto vira serviço gerenciado da Oracle.Of the six original components, four step aside. What stays: Object Storage — which speaks the S3 API and therefore remains neutral — and the Medallion model, which is data design rather than tooling. The rest becomes Oracle managed services.

01 / FonteSource

Oracle Database

Banco de dados transacional existente. A extração é feita via JDBC com queries incrementais por updated_at, minimizando impacto em produção.

Existing transactional database. Extraction via JDBC with incremental queries by updated_at, minimizing production impact.

JDBC Incremental CDC
01 / FonteSource

Oracle Database

Cenário A: migra para Autonomous Database ou Exadata Cloud, na mesma VCN do lake — a extração deixa de atravessar a WAN. Cenário B: permanece on-premise e a ingestão passa por FastConnect, igual à opção Azure.

Scenario A: moves to Autonomous Database or Exadata Cloud, in the same VCN as the lake — extraction no longer crosses the WAN. Scenario B: stays on-premise and ingestion goes over FastConnect, same as the Azure option.

Autonomous DB Exadata FastConnect
02 / IngestãoIngestion

OCI Data Integration

ETL gerenciado da Oracle, com editor visual em vez de DAGs em Python. Substitui Airflow e Spark de uma vez. Menos para operar — e nenhum pipeline que você possa levar embora.

Oracle's managed ETL, with a visual editor instead of Python DAGs. Replaces Airflow and Spark at once. Less to operate — and no pipeline you can take with you.

GerenciadoManaged Editor visualVisual editor GoldenGate
02 / OrquestraçãoOrchestration

Apache Airflow

Coração do ETL. Define, agenda e monitora cada passo do pipeline como um DAG. Interface web para gestão, backfill e alertas.

The ETL's heart. Defines, schedules, and monitors each pipeline step as a DAG. Web UI for management, backfill, and alerts.

DAGs Celery Backfill
03 / ProcessamentoProcessing

Apache Spark

Engine de transformações distribuídas. Executa dentro dos DAGs do Airflow para joins, deduplicação e aggregações em larga escala.

Distributed transformation engine. Runs inside Airflow DAGs for large-scale joins, deduplication, and aggregations.

PySpark Parquet SQL
0403 / ArmazenamentoStorage

MinIO

OCI Object Storage

Object storage S3-compatible on-premise. Armazena Bronze, Silver e Gold em Parquet. Também funciona como deep storage do Pinot.

On-premise S3-compatible object storage. Stores Bronze, Silver, and Gold in Parquet. Also serves as Pinot's deep storage.

O mesmo MinIO, agora em AKS sobre Azure Managed Disks (Premium SSD v2). Mantém a API S3 — é justamente ela que evita o lock-in: nenhum job conhece a Azure, todos falam S3.

The same MinIO, now on AKS over Azure Managed Disks (Premium SSD v2). It keeps the S3 API — which is precisely what avoids lock-in: no job knows about Azure, they all speak S3.

Aqui o MinIO simplesmente sai de cena: o Object Storage da OCI já fala API S3 nativa. Some o StatefulSet, os 120 TB de disco e a questão da AGPL — e mesmo assim este continua sendo o único componente portável da arquitetura.

Here MinIO simply leaves the picture: OCI Object Storage already speaks the native S3 API. The StatefulSet, the 120 TB of disk, and the AGPL question all disappear — and even so, this remains the only portable component in the architecture.

S3 API ZFS Managed Disks GerenciadoManaged IAM
04 / Analytics

Autonomous Data Warehouse

Data warehouse gerenciado, com tuning e escala automáticos. Cobre bem BI e relatórios, mas é um motor de warehouse, não de OLAP em tempo real: para o padrão de sub-100ms com alta concorrência do Pinot, precisa ser medido antes de prometer.

Managed data warehouse with automatic tuning and scaling. It covers BI and reporting well, but it is a warehouse engine, not a real-time OLAP one: matching Pinot's sub-100ms, high-concurrency profile has to be measured before it is promised.

SQL Auto-scaling A validarTo validate
05 / Analytics

Apache Pinot

OLAP distribuído de ultra-baixa latência. Indexa segmentos da camada Gold e responde queries em menos de 100ms, mesmo com bilhões de linhas.

Ultra-low latency distributed OLAP. Indexes Gold layer segments and answers queries in under 100ms, even with billions of rows.

StarTree SQL <100ms
05 / BI

Oracle Analytics Cloud

BI gerenciado, integrado ao ADW e ao Identity Domains para SSO e segurança por linha. Maduro e completo — mas os dashboards passam a viver num formato proprietário, e migrar depois significa refazê-los.

Managed BI, integrated with ADW and Identity Domains for SSO and row-level security. Mature and complete — but dashboards now live in a proprietary format, and migrating later means rebuilding them.

Dashboards SSO ProprietárioProprietary
06 / BI

Apache Superset

Plataforma de BI open source. Conecta-se nativamente ao Pinot via SQL. Dashboards interativos, exploração ad-hoc e alertas automáticos.

Open source BI platform. Connects natively to Pinot via SQL. Interactive dashboards, ad-hoc exploration, and automated alerts.

Dashboards SQL Lab Alertas
03

ETL com
Airflow
ETL with
Airflow

Cada pipeline é um DAG: um grafo de tarefas com dependências, retentativas e alertas configurados. Gestão completa via interface web. Each pipeline is a DAG: a task graph with dependencies, retries, and alerts configured. Full management via web UI.

DAG de ingestão diária: Oracle → MinIO → Pinot Daily ingestion DAG: Oracle → MinIO → Pinot
check_oracle _conn extract_tables OracleToMinIO validate_bronze GreatExpectations log_metadata PythonOperator spark_transform SparkSubmitOp write_gold MinIOOperator trigger_pinot HTTP Sensor schedule: @daily · retries: 3 · alertas: e-mail + Slack
ExtraçãoExtract

Oracle → Bronze

  • Conexão JDBC com pool gerenciado pelo Airflow
  • JDBC connection with Airflow-managed pool
  • Query incremental por coluna updated_at
  • Incremental query by updated_at column
  • Full load semanal para tabelas sem timestamp
  • Weekly full load for tables without timestamps
  • Escrita em JSON.gz no bucket Bronze do MinIO
  • Writes JSON.gz to MinIO Bronze bucket
  • Metadados: volume, checksum, timestamp de extração
  • Metadata: volume, checksum, extraction timestamp
TransformaçãoTransform

Bronze → Silver → Gold

  • Spark lê Bronze via API S3 do MinIO
  • Spark reads Bronze via MinIO S3 API
  • Silver: dedup, tipagem, limpeza de nulos
  • Silver: dedup, type casting, null cleanup
  • Particionamento por ano/mês/dia
  • Partitioned by year/month/day
  • Gold: joins entre domínios, KPIs de negócio
  • Gold: cross-domain joins, business KPIs
  • Saída em Parquet com compressão Snappy
  • Output in Parquet with Snappy compression
CargaLoad

Gold → Pinot

  • Airflow chama REST API do Pinot Controller
  • Airflow calls Pinot Controller REST API
  • Job offline: lê Parquet direto do MinIO
  • Offline job: reads Parquet directly from MinIO
  • Pinot persiste segmentos no bucket pinot-deep
  • Pinot persists segments in pinot-deep bucket
  • Dados disponíveis para query em menos de 5 min
  • Data available for querying in under 5 minutes
  • StarTree Index gerado automaticamente
  • StarTree Index generated automatically
03

ETL com
Data Integration
ETL with
Data Integration

Em vez de DAGs em Python versionados no Git, o pipeline é montado num editor visual gerenciado pela Oracle. É mais rápido de começar e bem mais difícil de versionar, revisar e migrar. Instead of Python DAGs versioned in Git, the pipeline is built in a visual editor managed by Oracle. Faster to start, and considerably harder to version, review, and migrate.

O que você troca, na prática: What you are trading, in practice: Some o Airflow, o Spark, o PostgreSQL de metadados, o Redis e todo o operacional que vinha junto — provavelmente as duas semanas mais trabalhosas do cronograma. Em troca, a lógica de transformação passa a viver dentro do serviço da Oracle: não há git diff de um pipeline, não há teste unitário de DAG, e recriar isso em outra nuvem significa refazer, não mover. Para times pequenos, costuma valer a pena; para times que já operam Airflow bem, raramente vale. Airflow, Spark, the metadata PostgreSQL, Redis, and all the operational work that came with them disappear — probably the two heaviest weeks of the schedule. In exchange, transformation logic now lives inside Oracle's service: there is no git diff of a pipeline, no unit test of a DAG, and recreating it on another cloud means rebuilding, not moving. For small teams this usually pays off; for teams already running Airflow well, it rarely does.
Alternativa intermediária: Middle-ground alternative: Nada impede manter Airflow e Spark no OKE e usar apenas o Object Storage e o ADW gerenciados. Você perde menos portabilidade e ainda assim elimina o MinIO e o Pinot da operação. Se a decisão de máximo gerenciado for por causa do tamanho da equipe, esse meio-termo merece ser avaliado antes. Nothing prevents keeping Airflow and Spark on OKE and using only managed Object Storage and ADW. You give up less portability and still remove MinIO and Pinot from operations. If the push toward fully managed is driven by team size, this middle ground deserves a look first.
04

Arquitetura
Medallion
Medallion
Architecture

O MinIO organiza os dados em três camadas, cada uma com propósito, formato e nível de qualidade específicos. Os dados nunca são sobrescritos: apenas promovidos. MinIO organizes data into three layers, each with a specific purpose, format, and quality level. Data is never overwritten: only promoted.

🥉
Bronze
Dados brutosRaw data
bronze/
  └─ oracle/
      ├─ clientes/
      │   └─ 2026/05/01/
      │       extract.json.gz
      ├─ pedidos/
      └─ estoque/
JSON / CSV Comprimido Retenção: 2 anosRetention: 2yr
🥈
Silver
Dados limposClean data
silver/
  └─ clientes/
      └─ year=2026/
          └─ month=05/
               part-000.parquet
Parquet Snappy ParticionadoPartitioned
🥇
Gold
Dados prontosBusiness-ready
gold/
  ├─ vendas_diarias/
  ├─ clientes_ativos/
  ├─ estoque_critico/
  └─ faturamento_mensal/
Parquet Pinot-ready Por domínioBy domain
05

Os três
servidores
The three
servers
Os três
node pools
The three
node pools

Cada servidor tem uma responsabilidade única. O isolamento facilita escalonamento independente e aumenta resiliência do conjunto.Each server has a unique responsibility. Isolation enables independent scaling and increases overall resilience. A mesma separação de responsabilidades, agora como node pools dedicados num único cluster AKS. A diferença prática: cada pool escala sozinho, sob demanda, e você paga só pelo que está ligado.The same separation of concerns, now as dedicated node pools in a single AKS cluster. The practical difference: each pool scales on its own, on demand, and you pay only for what is running. Esta é a seção mais curta das três — e é exatamente esse o argumento. Não há servidor para dimensionar nem node pool para ajustar: os serviços são gerenciados e escalam sozinhos. O que sobra de infraestrutura própria é um cluster OKE pequeno, e mesmo ele é opcional.This is the shortest of the three versions of this section — and that is precisely the argument. There is no server to size and no node pool to tune: the services are managed and scale on their own. What remains of your own infrastructure is one small OKE cluster, and even that is optional.

🗄️ Oracle DB Existente JDBC SRV-01 · ETL ⚙️ Apache Airflow 2.9 Webserver · Scheduler · Celery 🔥 Apache Spark 3.5 Master · Workers locais 🐘 PostgreSQL Airflow meta 🚀 Redis Celery broker 10.0.0.11 S3 API SRV-02 · Storage 🪣 MinIO API :9000 · Console :9001 📂 bronze · silver · gold · pinot-deep Buckets + IAM Policies 💾 ZFS · RAID-Z2 12× 10TB HDD + 4× 2TB NVMe cache 10.0.0.12 S3 API SRV-03 · Analytics ⚡ Apache Pinot 1.2 Controller · Broker · Server 📊 Apache Superset BI · Dashboards · Alertas 🐘 ZooKeeper Pinot coord. 🔐 Nginx TLS + proxy 10.0.0.13 👥 Analistas VLAN 10.0.0.0/24 · Switch 10GbE
🗄️ Oracle DB On-prem JDBC VPN / ExpressRoute AZURE VNET · PRIVATE SUBNET AKS CLUSTER pool-etl ⚙️ Airflow 2.9 KubernetesExecutor 🔥 Spark 3.5 Spark on K8s 🐘 Azure DB Postgres Airflow meta (PaaS) D8s v5 · autoscale 2→8 S3 pool-storage 🪣 MinIO StatefulSet · 4 réplicas 📂 bronze · silver · gold pinot-deep 💾 Managed Disks Premium SSD v2 · 120 TB E8s v5 · fixo 4 S3 pool-analytics ⚡ Pinot 1.2 Ctrl · Broker · Server 📊 Superset BI · Dashboards 🐘 ZooKeeper Pinot coord. L16s v3 · NVMe local 🔐 App Gateway WAF · TLS · Entra ID 👥 Analistas Analysts Key Vault · Entra ID · Azure Monitor · Container Registry · Backup Region: brazilsouth Tudo em subnet privada · sem IP público nos pods · egress via NAT Gateway All in a private subnet · no public IP on pods · egress via NAT Gateway
01
Servidor ETL ETL Server
Airflow + Spark · 10.0.0.11
CPU 2× Xeon Silver 4314
32 cores / 64 threads
RAM 128 GB DDR4 ECC
(expansível a 256 GB) (expandable to 256 GB)
OS Disk 2× 500 GB NVMe RAID-1
Disco tempTemp disk 4× 2 TB NVMe
(Spark shuffle) (Spark shuffle)
Rede 2× 10GbE bonding
Airflow 2.9 Spark 3.5 PostgreSQL 16 Redis 7 Java 17 Python 3.12 Docker
02
Servidor Storage Storage Server
MinIO · 10.0.0.12
CPU 1× Xeon Silver 4310
12 cores / 24 threads
RAM 64 GB DDR4 ECC
(MinIO é I/O-bound) (MinIO is I/O-bound)
OS Disk 2× 500 GB SSD RAID-1
Data Lake 12× 10 TB HDD
120 TB bruto · RAID-Z2
Cache 4× 2 TB NVMe
ZFS L2ARC
Rede 2× 10GbE + 1GbE gestãomgmt
MinIO AGPL ZFS on Linux Ubuntu 24.04 Prometheus
03
Servidor Analytics Analytics Server
Pinot + Superset · 10.0.0.13
CPU 2× Xeon Gold 6338
64 cores / 128 threads
RAM 256 GB DDR4 ECC
(Pinot é memory-intensive) (Pinot is memory-intensive)
OS Disk 2× 500 GB NVMe RAID-1
SegmentosSegments 8× 4 TB NVMe
32 TB · acesso sub-mssub-ms access
Rede 2× 10GbE + 1GbE gestãomgmt
Pinot 1.2 Superset 4.x ZooKeeper 3.9 Nginx Java 21 Docker
01
Pool de ETL ETL pool
Airflow + Spark · pool-etl
VM SKU Standard_D8s_v5
8 vCPU / 32 GB
EscalaScale 2 → 8 nós (autoscale) 2 → 8 nodes (autoscale)
picos só na janela do ETL bursts only during the ETL window
Disco tempTemp disk NVMe efêmero do nó Node ephemeral NVMe
(Spark shuffle)
MetadadosMetadata Azure Database for PostgreSQL
(gerenciado · substitui o PG local) (managed · replaces local PG)
RedeNetwork Subnet privada · Azure CNI Private subnet · Azure CNI
Airflow 2.9 Spark 3.5 KubernetesExecutor Helm Java 17 Python 3.12
02
Pool de Storage Storage pool
MinIO · pool-storage
VM SKU Standard_E8s_v5
8 vCPU / 64 GB
EscalaScale 4 nós fixos (StatefulSet) 4 fixed nodes (StatefulSet)
erasure coding do MinIO MinIO erasure coding
Data Lake 120 TB · Premium SSD v2 120 TB · Premium SSD v2
(IOPS e throughput ajustáveis à parte) (IOPS and throughput tuned separately)
Camada friaCold tier Blob Archive p/ Bronze > 90 dias Blob Archive for Bronze > 90 days
Backup Snapshots de disco + Azure Backup Disk snapshots + Azure Backup
MinIO AGPL Managed Disks Blob Archive Azure Backup
03
Pool de Analytics Analytics pool
Pinot + Superset · pool-analytics
VM SKU Standard_L16s_v3
16 vCPU / 128 GB
EscalaScale 2 → 4 nós (autoscale) 2 → 4 nodes (autoscale)
acompanha a concorrência de queries follows query concurrency
SegmentosSegments 2× 1,92 TB NVMe local (~3,8 TB/nó) 2× 1.92 TB local NVMe (~3.8 TB/node)
acesso sub-ms · é o que sustenta o <100ms sub-ms access · this is what sustains the <100ms
Deep storageDeep storage MinIO (pool-storage) via S3 MinIO (pool-storage) over S3
EntradaIngress App Gateway + WAF · SSO via Entra ID App Gateway + WAF · SSO via Entra ID
Pinot 1.2 Superset 4.x ZooKeeper 3.9 App Gateway Entra ID Java 21
01
Ingestão + Storage Ingestion + Storage
Data Integration · Object Storage
ETL OCI Data Integration
(gerenciado · sem servidor) (managed · serverless)
Data Lake Object Storage
Standard + Archive · API S3 Standard + Archive · S3 API
EscalaScale Elástica · sem capacidade a provisionar Elastic · no capacity to provision
CDC GoldenGate (opcional, se precisar de tempo real) (optional, if real time is needed)
Data Integration Object Storage S3 API GoldenGate
02
Analytics + BI Analytics + BI
ADW · Analytics Cloud
OLAP Autonomous Data Warehouse
auto-tuning · auto-scaling auto-tuning · auto-scaling
BI Oracle Analytics Cloud
dashboards · alertas · RLS dashboards · alerts · RLS
LatênciaLatency a medir na POC — ADW é warehouse, não OLAP de tempo real to be measured in the POC — ADW is a warehouse, not real-time OLAP
IdentidadeIdentity OCI Identity Domains (SSO)
ADW Analytics Cloud Identity Domains
03
O que sobra seu What stays yours
OKE · opcionaloptional
OKE Cluster pequeno para jobs próprios, scripts e ferramentas que você não quer entregar ao fornecedor. A small cluster for your own jobs, scripts, and tooling you would rather not hand to the vendor.
PortávelPortable Object Storage (API S3) e o modelo Medallion — só isso atravessa uma migração futura sem reescrita. Object Storage (S3 API) and the Medallion model — only these survive a future migration without a rewrite.
PresoLocked in Pipelines, modelo do warehouse e dashboards. Refazer, não mover. Pipelines, warehouse model, and dashboards. Rebuild, not move.
OKE Terraform S3 API

Portas & rede Ports & network

Acesso & rede Access & network

ServiçoService ServidorServer PortaPort AcessoAccess
Airflow UISRV-018080LAN internaInternal LAN
Spark UISRV-014040Equipe dadosData team
MinIO APISRV-029000SRV-01 e SRV-03SRV-01 and SRV-03
MinIO ConsoleSRV-029001AdminsAdmins
Pinot ControllerSRV-039000SRV-01 (Airflow)SRV-01 (Airflow)
Pinot Broker SQLSRV-038099Superset + devsSuperset + devs
Apache SupersetSRV-03443 (Nginx)Todos usuários · HTTPSAll users · HTTPS
ZooKeeperSRV-032181Interno SRV-03SRV-03 internal
ServiçoService Endereço no clusterIn-cluster address PortaPort AcessoAccess
Airflow UIairflow-web.etl8080App Gateway · Entra IDApp Gateway · Entra ID
Spark UIspark-driver.etl4040Port-forward · equipe dadosPort-forward · data team
MinIO APIminio.storage9000Interno ao clusterCluster-internal
MinIO Consoleminio-console.storage9001Admins · via App GatewayAdmins · via App Gateway
Pinot Controllerpinot-controller.analytics9000Airflow (pool-etl)Airflow (pool-etl)
Pinot Broker SQLpinot-broker.analytics8099Superset + devsSuperset + devs
Apache Supersetsuperset.analytics443 (App Gateway)(App Gateway)Todos usuários · HTTPS + SSOAll users · HTTPS + SSO
ZooKeeperzookeeper.analytics2181Interno ao namespaceNamespace-internal
ServiçoService TipoType AcessoAccess
Object StorageGerenciadoManagedEndpoint S3 · private endpoint na VCNS3 endpoint · private endpoint in the VCN
Data IntegrationGerenciadoManagedConsole OCI · Identity DomainsOCI Console · Identity Domains
Autonomous DWGerenciadoManagedSQL · private endpoint · mTLSSQL · private endpoint · mTLS
Analytics CloudGerenciadoManagedHTTPS · SSO · todos usuáriosHTTPS · SSO · all users
Autonomous DB (cenário A)(scenario A)GerenciadoManagedMesma VCN · sem WANSame VCN · no WAN
Oracle on-prem (cenário B)(scenario B)SeuYoursFastConnect ou VPNFastConnect or VPN
OKE (opcional)(optional)SeuYoursSubnet privadaPrivate subnet
Menos portas, mais contrato. Fewer ports, more contract. Repare que a tabela encolheu: quase não há porta para liberar nem serviço para expor, porque a Oracle opera tudo por trás de endpoints privados e do Identity Domains. É uma superfície de ataque menor e menos trabalho de rede — em troca, o que antes era configuração sua vira cláusula de contrato e SLA. Notice the table shrank: there is barely a port to open or a service to expose, because Oracle runs everything behind private endpoints and Identity Domains. That is a smaller attack surface and less network work — in exchange, what used to be your configuration becomes a contract clause and an SLA.
Nenhum IP público. No public IPs. Os pods vivem numa subnet privada; a única porta de entrada é o Application Gateway com WAF, autenticando contra o Entra ID. Saída para a internet só via NAT Gateway, com IP fixo — o que facilita liberar a origem no firewall do Oracle on-premise. Pods live in a private subnet; the only way in is the Application Gateway with WAF, authenticating against Entra ID. Egress goes through a NAT Gateway with a fixed IP — which makes allow-listing the source on the on-premise Oracle firewall straightforward.
06

Roadmap de implantação Deployment
roadmap

Infraestrutura na semana 3. ETL rodando e engenheiros com acesso ao MinIO na semana 7. Pinot com queries abaixo de 100ms na semana 10. Superset, monitoramento e handoff na semana 13.Infrastructure by week 3. ETL running and engineers with MinIO access by week 7. Pinot with sub-100ms queries by week 10. Superset, monitoring, and handoff by week 13. Cinco semanas a menos: não há compra, entrega nem racking de hardware. Infraestrutura na semana 1. ETL rodando na semana 4. Pinot abaixo de 100ms na semana 6. Superset, monitoramento e handoff na semana 8.Five weeks shorter: there is no hardware to buy, ship, or rack. Infrastructure by week 1. ETL running by week 4. Pinot under 100ms by week 6. Superset, monitoring, and handoff by week 8. O mais curto dos três, porque quase não há o que instalar. A maior parte do tempo vai para modelagem, migração do Oracle (cenário A) e reconstrução dos dashboards no OAC — não para infraestrutura.The shortest of the three, because there is almost nothing to install. Most of the time goes into modeling, migrating Oracle (scenario A), and rebuilding dashboards in OAC — not into infrastructure.

Fase 1 · Semanas 1–3Phase 1 · Weeks 1–3 Fase 1 · Semana 1Phase 1 · Week 1 Fase 1 · Semana 1Phase 1 · Week 1

Infraestrutura base Base infrastructure

Provisionamento dos 3 servidores, instalação do SO, configuração de rede/VLAN, instalação do MinIO com ZFS. Definição de políticas IAM e criação dos buckets Bronze, Silver, Gold e pinot-deep.

Provision 3 servers, OS installation, network/VLAN config, MinIO with ZFS. Define IAM policies and create Bronze, Silver, Gold, and pinot-deep buckets.

Terraform provisiona VNet, subnets, cluster AKS e os 3 node pools. Deploy do MinIO via Helm sobre Managed Disks, criação dos buckets Bronze, Silver, Gold e pinot-deep, e das políticas IAM. Conectividade com o Oracle on-premise (VPN ou ExpressRoute) validada.

Terraform provisions the VNet, subnets, AKS cluster, and the 3 node pools. MinIO deployed via Helm over Managed Disks, Bronze/Silver/Gold/pinot-deep buckets and IAM policies created. Connectivity to the on-premise Oracle (VPN or ExpressRoute) validated.

Terraform provisiona a VCN, subnets e os buckets do Object Storage, com políticas IAM e Identity Domains. Decisão do cenário A ou B do Oracle e, no caso do B, FastConnect validado. Não há cluster nem storage para instalar.

Terraform provisions the VCN, subnets, and Object Storage buckets, with IAM policies and Identity Domains. Decision between Oracle scenario A or B and, for B, FastConnect validated. There is no cluster or storage to install.

Fase 2 · Semanas 4–7Phase 2 · Weeks 4–7 Fase 2 · Semanas 2–4Phase 2 · Weeks 2–4 Fase 2 · Semanas 2–3Phase 2 · Weeks 2–3

ETL Oracle → MinIO Bronze/Silver ETL Oracle → MinIO Bronze/Silver

Instalação do Airflow e Spark, configuração das conexões Oracle, desenvolvimento dos DAGs de extração incremental, transformações Bronze→Silver. Validação de qualidade com Great Expectations.

Install Airflow and Spark, configure Oracle connections, develop incremental extraction DAGs, Bronze→Silver transformations. Quality validation with Great Expectations.

Migração do Oracle para Autonomous DB / Exadata, se for o cenário A — é a tarefa mais pesada desta fase e precisa de janela combinada. Pipelines de extração e transformação Bronze→Silver montados no Data Integration.

Oracle migration to Autonomous DB / Exadata, if scenario A — the heaviest task in this phase, requiring an agreed window. Extraction and Bronze→Silver transformation pipelines built in Data Integration.

Fase 3 · Semanas 8–10Phase 3 · Weeks 8–10 Fase 3 · Semanas 5–6Phase 3 · Weeks 5–6 Fase 3 · Semana 4Phase 3 · Week 4

Camada Gold + Apache Pinot Gold layer + Apache Pinot

Modelagem dos dados de negócio (Gold), transformações Spark, instalação do Pinot e ZooKeeper, criação dos schemas Pinot, pipeline de ingestão Gold→Pinot via REST API.

Business data modeling (Gold), Spark transformations, Pinot and ZooKeeper install, Pinot schema creation, Gold→Pinot ingestion pipeline via REST API.

Modelagem da camada Gold e carga no Autonomous Data Warehouse. É aqui que a latência precisa ser medida de verdade, com dados e concorrência reais — antes de prometer qualquer número ao usuário final.

Gold layer modeling and load into Autonomous Data Warehouse. This is where latency has to be genuinely measured, with real data and real concurrency — before promising any number to end users.

Fase 4 · Semanas 11–13Phase 4 · Weeks 11–13 Fase 4 · Semanas 7–8Phase 4 · Weeks 7–8 Fase 4 · Semanas 5–6Phase 4 · Weeks 5–6

Superset + monitoramento + entrega Superset + monitoring + delivery

Instalação do Superset, conexão com Pinot, primeiros dashboards, configuração de usuários e row-level security, alertas, Prometheus + Grafana para monitoramento dos servidores. Handoff para a equipe.

Superset install, Pinot connection, first dashboards, users and row-level security setup, alerts, Prometheus + Grafana for server monitoring. Team handoff.

Instalação do Superset, conexão com Pinot, primeiros dashboards, SSO via Entra ID e row-level security, alertas. Monitoramento com Azure Monitor + Managed Grafana. Terraform e Helm charts entregues no seu repositório — é o que garante que a saída da Azure continue sendo uma opção real.

Superset install, Pinot connection, first dashboards, SSO via Entra ID and row-level security, alerts. Monitoring with Azure Monitor + Managed Grafana. Terraform and Helm charts handed over in your own repository — which is what keeps leaving Azure a real option.

Construção dos dashboards no Oracle Analytics Cloud — não é migração, é reconstrução, já que não há importação do Superset. SSO e segurança por linha via Identity Domains, alertas e handoff.

Building dashboards in Oracle Analytics Cloud — not a migration but a rebuild, since there is no Superset import path. SSO and row-level security via Identity Domains, alerts, and handoff.

Custo de licença: R$ 0. License cost: R$ 0. Airflow, Spark, MinIO, Pinot e Superset são 100% open source. O único custo é o hardware on-premise (~USD 58.500 para médio porte) e o time de implantação. Airflow, Spark, MinIO, Pinot, and Superset are 100% open source. The only cost is on-premise hardware (~USD 58,500 for mid-size) and the deployment team.
Custo de licença: R$ 0. O que muda é a forma de pagar o resto. License cost: R$ 0. What changes is how you pay for the rest. O stack continua 100% open source — nenhuma licença entra na conta. A diferença é que o hardware deixa de ser CAPEX (~USD 58.500 de uma vez) e vira OPEX mensal. Como ordem de grandeza, com preço de tabela e os SKUs desta proposta, o cluster fica na casa de USD 6–9 mil/mês; com reserved instances de 3 anos, tipicamente 40–60% menos. Ou seja: a Azure ganha no primeiro ano e passa a perder por volta do segundo ou terceiro — o ponto exato depende do seu volume e da negociação com a Microsoft, e deve ser calculado antes de decidir. The stack is still 100% open source — no license enters the equation. The difference is that hardware stops being CAPEX (~USD 58,500 up front) and becomes monthly OPEX. As an order of magnitude, at list price with the SKUs in this proposal, the cluster lands around USD 6–9k/month; with 3-year reserved instances, typically 40–60% less. In other words: Azure wins in year one and starts losing somewhere around year two or three — the exact crossover depends on your volume and your negotiation with Microsoft, and should be calculated before deciding.
Duas ressalvas técnicas, ditas na frente. Two technical caveats, stated up front. (1) O NVMe local da série L é efêmero: se o nó for desalocado, os segmentos daquele nó somem. Não é um problema — é exatamente por isso que o Pinot mantém o deep storage no MinIO e recarrega sozinho — mas significa que o pool de analytics nunca deve ser tratado como fonte de verdade. (2) Se o Oracle continuar on-premise, a extração passa a atravessar a WAN todo dia; com volume alto, ExpressRoute deixa de ser luxo e vira requisito, e o custo dele precisa entrar na conta acima. (1) The L-series local NVMe is ephemeral: if a node is deallocated, that node's segments are gone. This is not a flaw — it is precisely why Pinot keeps deep storage on MinIO and reloads on its own — but it means the analytics pool must never be treated as a source of truth. (2) If Oracle stays on-premise, extraction now crosses the WAN every day; at high volume ExpressRoute stops being a luxury and becomes a requirement, and its cost has to go into the figure above.
Aqui não dá para dizer "R$ 0 em licenças". Here you cannot claim "R$ 0 in licenses". Esta é a diferença comercial mais importante entre as três abas. ADW e Oracle Analytics Cloud são produtos licenciados: o custo deixa de ser só infraestrutura e passa a incluir software, sob contrato com a Oracle. Não coloco um número aqui de propósito — diferente de VM, esse preço é negociado, e depende do seu contrato atual, do volume e de créditos de migração, que costumam ser agressivos para quem leva banco Oracle para a OCI. Peça a proposta antes de comparar com os números das outras duas abas, senão a comparação não é honesta. This is the most important commercial difference between the three tabs. ADW and Oracle Analytics Cloud are licensed products: cost stops being infrastructure alone and starts including software, under contract with Oracle. I deliberately do not put a number here — unlike a VM, this price is negotiated, and depends on your current agreement, your volume, and migration credits, which tend to be aggressive for customers bringing an Oracle database to OCI. Get the quote before comparing with the figures in the other two tabs, otherwise the comparison is not an honest one.
Três ressalvas técnicas, ditas na frente. Three technical caveats, stated up front. (1) O <100ms das outras abas vem do Pinot, que é OLAP de tempo real. O ADW é um data warehouse: excelente para BI e relatório, mas o perfil de latência sob alta concorrência é outro. Trate esse número como a validar, não como herdado. (2) Os dashboards do Superset não migram para o OAC — são refeitos. (3) Concentrar banco, ETL, warehouse e BI num fornecedor único enfraquece sua posição na renovação. É uma decisão comercial legítima, mas deve ser tomada com os olhos abertos. (1) The <100ms in the other tabs comes from Pinot, a real-time OLAP engine. ADW is a data warehouse: excellent for BI and reporting, but its latency profile under high concurrency is a different thing. Treat that number as to be validated, not inherited. (2) Superset dashboards do not migrate to OAC — they are rebuilt. (3) Concentrating database, ETL, warehouse, and BI in a single vendor weakens your position at renewal. That is a legitimate commercial decision, but it should be made with eyes open.