EMR AI Assistant is an intelligent operations and development service provided by Alibaba Cloud EMR. This topic describes the core advantages, functional modules, coverage scenarios, and plans of AI Assistant.
Overview
EMR AI Assistant is an intelligent operations and development service provided by Alibaba Cloud EMR. Powered by large language models, AI Assistant connects to your EMR clusters and instances, understands their status and monitoring data, and delivers full-stack AI-driven operations capabilities — from Q&A consulting and performance diagnostics to health inspections and proactive alerts — serving as your 24/7 big data expert.
Core advantages
Connected to real clusters, not generic Q&A
AI Assistant connects directly to your StarRocks instances and reads runtime data such as system tables, query profiles, and monitoring metrics. All diagnostics and recommendations are based on the actual state of your cluster, not generic knowledge.
From reactive responses to proactive operations
AI Assistant does not only respond when you ask. By configuring daily reports, inspection reports, and alert notifications, AI Assistant continuously monitors cluster health around the clock and proactively pushes findings to DingTalk, Lark, or WeCom — no manual monitoring required.
Multiple access points, always available
Access AI Assistant through the EMR console, DingTalk/Lark/WeCom IM channels, or API. It integrates into your existing operations workflow without requiring you to switch tools.
Functional modules
AI Assistant includes the following core modules:
Chat window
A built-in AI chat interface in the console that supports multi-turn conversations with AI Assistant.
Supports natural language queries. AI Assistant automatically identifies intent and connects to your cluster for analysis.
Supports contextual multi-turn conversations, allowing you to progressively drill down into issues within a single session.
Conversation history is automatically saved. You can review previous diagnostic records at any time.
Supports selecting associated instances and expert skills to analyze different instances for different scenarios.
Task center
Manages scheduled tasks and historical execution records of AI Assistant.
Scheduled tasks:
Configure the push schedule and frequency for daily reports.
Configure the execution cycle for inspection reports.
Configure thresholds and push rules for slow SQL alerts.
Supports enabling, pausing, and deleting tasks.
Execution records:
View historical execution results for all tasks.
View the complete content of each daily report, inspection, or alert.
Filter records by time, type, or status.
View failure reasons when execution fails.
IM configuration
Integrate AI Assistant with enterprise IM tools so that team members can use it directly from work group chats.
Supports DingTalk, Lark, and WeCom.
Configure the Webhook URL and signing key for the corresponding group.
API/SDK integration
Embed EMR AI Assistant capabilities into your enterprise platform and automated operations workflows through APIs and SDKs. For more information, see Call the EMR AI Assistant API.
Coverage scenarios
Performance diagnostics
When your cluster experiences latency spikes, query timeouts, or resource saturation, AI Assistant quickly identifies the root cause and provides remediation steps.
Scenario | What AI Assistant does | Example |
Slow query identification | Analyzes top SQL statements, sorted by duration, CPU usage, and rows scanned | "Any slow queries recently?" |
Execution plan analysis | Deep-dives into query profiles to pinpoint operator-level bottlenecks | "This query took 30 seconds — help me check the profile" |
Resource bottleneck diagnosis | Analyzes CPU, memory, disk, and I/O usage to identify hotspots | "CPU suddenly spiked to 100%" |
Compaction backlog | Analyzes version accumulation and write amplification, then provides tuning recommendations | "Version count alert triggered — what should I do?" |
Cluster operations
Covers routine inspections, auto scaling, configuration management, and other day-to-day operations tasks.
Scenario | What AI Assistant does | Example |
Health inspection | Performs a comprehensive 8-dimension health check, outputs a health score and optimization recommendations | "Run a full health check" |
Auto scaling | Analyzes historical workloads to assess whether scaling up or down is needed | "Do we need to add nodes before peak traffic?" |
Connection diagnostics | Troubleshoots connection failures, timeouts, and IP allowlist issues | "Applications suddenly cannot connect to the cluster" |
Configuration tuning | Compares against best practices and recommends parameter adjustments | "How many compaction threads should I set?" |
Instance management | Views instance information, version details, account permissions, and more | "What are the specs of my current instance?" |
Troubleshooting
Scenario | What AI Assistant does | Example |
Error diagnosis | Matches against known issue databases and provides remediation steps | "I got this error — how do I fix it?" |
Node failure | Analyzes BE/FE crash causes and provides recovery steps | "A BE node is down — what happened?" |
OOM analysis | Identifies memory consumption sources and provides parameter tuning recommendations | "My query hit an OOM error" |
Data load failure | Analyzes data load error causes and provides remediation steps | "Stream Load returned an error" |
Proactive operations
AI Assistant can automatically execute analysis tasks on a schedule and push results to IM channels.
Push type | Trigger | Content |
Daily report | Daily schedule | Cluster health score + anomaly event diagnostics + trend comparison + optimization recommendations |
Inspection report | Weekly schedule | 8-dimension health score + comprehensive check + optimization checklist |
Alert notification | Monitoring event | Alert details + AI root cause analysis + recommended actions |
Slow SQL alert | Threshold trigger | Slow SQL details + profile analysis + optimized SQL |
All scheduled tasks and push records can be viewed and managed in Task Center.
Consulting
Scenario | What AI Assistant does | Example |
Product usage | Answers usage questions based on the latest StarRocks documentation | "How do I create a materialized view?" |
Best practices | Provides best practices for table design, data loading, and query optimization | "How do I choose the number of buckets?" |
Feature consultation | Explains StarRocks features and applicable scenarios | "When should I use Colocate Join?" |
Ecosystem integration | Guidance for integrating external data sources such as Paimon, Iceberg, Fluss, and Kafka | "How do I configure a Paimon catalog?" |
Development assistance
Scenario | What AI Assistant does | Example |
Table design | Recommends partitioning and bucketing strategies, sort keys, and index configurations based on query patterns | "How should I design a table for 500 million rows per day?" |
SQL optimization | SQL rewriting, join strategy adjustments, and predicate pushdown verification | "Help me optimize this SQL query" |
Data loading | Generates Stream Load, Routine Load, and Broker Load configurations | "How do I load Kafka data in real time?" |
Materialized views | Creation recommendations, refresh strategy configuration, status checks, and anomaly resolution | "Can I accelerate this report query with a materialized view?" |
Plans and specifications
Each user receives a free token quota of 1 million tokens. After the free quota is exhausted, you can choose from the following plans based on your usage:
Specification | Free | Basic | Professional | Enterprise |
Token quota | 1 million | 6 million | 125 million | 315 million |
Console chat window | Supported | Supported | Supported | Supported |
IM channels (DingTalk/Lark/WeCom) | Supported | Supported | Supported | Supported |
Expert skills | Supported | Supported | Supported | Supported |
Task center | Supported | Supported | Supported | Supported |
API | 5 calls/day | 5 calls/day | 100 calls/day | Unlimited |
Proactive alert push, slow SQL alerts | Supported | Supported | Supported | Supported |
Concurrency | 1 | 1 | 5 | 10 |
For a quick start guide, see EMR AI Assistant quick start.
Disclaimer
The output of this service is generated by an AI model. We cannot fully guarantee the safety, reliability, availability, and continuous stability of the service or the compliance, completeness, and accuracy of the generated content. The generated content does not represent the positions or views of Alibaba Cloud. We will continuously improve service quality but do not guarantee the availability or reliability of the service and are not responsible for the results of your use of this service. Exercise caution when evaluating the generated content and do not rely on it excessively. You are fully responsible for any judgments or actions taken based on the generated content that result in any loss or damage to you, other users, or Alibaba Cloud.
You are solely responsible for all your usage behavior. Ensure that any content you publish, upload, link, or provide through the service is lawful and compliant, does not harm public order, does not infringe upon the legitimate rights of others, and does not fabricate or disseminate false information.
Using the diagnostic features of this service requires the collection of data related to the diagnosed instance, including basic information about your database instance, system table contents, instance monitoring metrics, and key error information from events and logs.
FAQ
Q: Does AI Assistant read my business data?
A: AI Assistant only reads system tables (such as information_schema), query profiles, monitoring metrics, and other operational data from your cluster for diagnostics and analysis. It does not read the content of your business tables.
Q: How are the permissions of AI Assistant controlled?
A: The scope of AI Assistant operations is controlled by RAM permissions. Currently, it primarily performs read-only diagnostics and analysis. Write operations (such as configuration changes) require your explicit confirmation in the conversation before execution.
Q: What happens when the token quota is exhausted?
A: When the token quota is exhausted, AI Assistant pauses the service. You can enable overage billing to continue using the service. We recommend monitoring your quota usage in the console and adjusting as needed.
Q: How do I manage multiple instances?
A: A single AI Assistant subscription can be associated with multiple instances under the same account. Specify the instance name in the conversation to switch the analysis target. Scheduled tasks can be configured with independent push rules for each instance.
Q: What should I do if a scheduled task fails?
A: You can view the failure reason in the execution records of Task Center. Common causes include: the instance is stopped, the IM channel connection is broken, or the token quota is insufficient. After the issue is resolved, the task automatically resumes at the next scheduled trigger.