diff --git a/skills/analytics-opensearch-expertise/.skilleval.yaml b/skills/analytics-opensearch-expertise/.skilleval.yaml new file mode 100644 index 0000000..686a9c7 --- /dev/null +++ b/skills/analytics-opensearch-expertise/.skilleval.yaml @@ -0,0 +1,3 @@ +audit: + ignore: + - STR-016 # README alongside SKILL.md is intentional diff --git a/skills/analytics-opensearch-expertise/CHANGELOG.md b/skills/analytics-opensearch-expertise/CHANGELOG.md new file mode 100644 index 0000000..b2623ec --- /dev/null +++ b/skills/analytics-opensearch-expertise/CHANGELOG.md @@ -0,0 +1,82 @@ +## [2.6.0] - 2026-07-24 +### Changed +- Remediation Reference moved to `references/remediation-reference.md`, loaded on demand via + `read_skill_resource` (mandatory Execution Flow step) — keeps SKILL.md compact so late-file + instructions are never truncated at skill ingestion. +### Added +- API compatibility guidance: the DevOps Agent tooling exposes the legacy + Elasticsearch-compatible client; legacy operation names listed first with modern fallbacks. +- CloudWatch fallback for DescribeDomainHealth (Shards.active / Shards.unassigned / + ClusterStatus.*) so shard checks are never skipped; namespace and dimensions stated + explicitly (AWS/ES, DomainName + ClientId). +- Category roll-up rule for the executive summary (worst finding wins). +- Instance RAM reference table for heap estimation. +- Explicit SKIPPED paths for 2.4 (metric unavailable) and 5.3 (AccessDenied); 3.7 + currently-yellow reporting clarified. + +# Changelog + +All notable changes to the `analytics-opensearch-expertise` skill are documented here. + +## [2.5.0] - 2026-07-24 +### Added +- Severity-override prohibition: verdict tiers come only from the matching Logic line, never + from general judgment; inline "do not escalate" guards on checks 4.1, 4.2, and 4.6. +- Check-ID fidelity rule: canonical 24-entry ID-to-name list embedded in the coverage rule; + renumbering, splitting, merging, or inventing checks is prohibited. + +## [2.4.0] - 2026-07-23 +### Changed +- Security severity rebalance: 4.1 (encryption at rest), 4.2 (node-to-node encryption), + 4.3 (HTTPS), 4.5 (access policy), and 4.6 (FGAC) are Warnings when a control is disabled + in isolation; 4.4 compound exposure (public + HTTPS off + FGAC off) is the only security + Critical. Criticals are reserved for compounded exposure and imminent operational risk. +- 3.3/3.4 latency checks demoted to informational-only: the millisecond markers are + heuristics with no AWS-documented threshold; latency is workload-dependent. +- 3.7 renamed "Cluster Status Duration" to "Cluster Status History". +- 1.4 version counting rule: "releases behind" is the count of later same-major versions + returned by ListVersions, never the numeric delta between version numbers. +- 1.4 fallback provenance note now lives inside the finding only, phrased for the reader; + API error names never appear in the report. + +## [2.3.0] - 2026-07-21 +### Added +- Remediation Reference: a 24-entry verbatim dictionary (why-it-matters, resolution steps, + and verified official AWS documentation links). Every non-passing finding pulls its entry + verbatim; documentation links come only from the dictionary, eliminating URL hallucination. +### Changed +- 4.4 Network Exposure escalation logic changed from OR to AND: Critical only when public + endpoint, HTTPS off, and FGAC off are all true; a single weak control remains a Warning + at 4.4 (it is still flagged by its own check). +- 1.4 upgrade candidates render as a set of direct single-hop targets, never a sequential + chain; upgrade recommendations propose a single hop to the latest same-major version. + +## [2.2.0] - 2026-07-21 +### Fixed +- 2.4 free-storage formula: the FreeStorageSpace Minimum statistic is per-node; the + percentage is computed against one node's volume (the previous formula understated free + space by the node count). +- 3.1 pinned to the Maximum statistic; Average is reported as context only. +### Added +- 1.4 fallback: derive upgrade candidates via ListVersions when GetCompatibleVersions is + unavailable; skip only if both APIs fail. +- Deterministic tier boundaries for 2.2 (shard density bands) and 3.6 (tier determined + only by 5xx count). +- Target Resolution rules: exact match proceeds; no target triggers discovery with options; + a near-miss target stops for explicit confirmation and is never silently substituted. +- Coverage rule hardening: no tier rounding by judgment; CloudWatch rows name the statistic + used; mandatory 24-row matrix self-count. + +## [2.1.0] - 2026-07-14 +### Changed +- 2.2 shard density computes the best-practice limit from the actual instance heap + (heap = min(32, RAM/2) GiB x 25 shards) instead of a fixed assumption; Critical tier + added at 4x the limit. +- 5.1 right-sizing rule made explicitly conjunctive (CPU < 20% AND JVM < 50%); suppressed + when JVM pressure is at or above 50%. + +## [1.0.0] - 2026-07-11 +### Added +- Initial release: 24 read-only control-plane checks across cluster health, storage and + shard strategy, performance, security posture, and cost optimization, with a structured + findings report, coverage matrix, and prioritized recommendations. diff --git a/skills/analytics-opensearch-expertise/README.md b/skills/analytics-opensearch-expertise/README.md new file mode 100644 index 0000000..1685151 --- /dev/null +++ b/skills/analytics-opensearch-expertise/README.md @@ -0,0 +1,53 @@ +# analytics-opensearch-expertise + +A read-only Amazon OpenSearch Service **domain health assessment** skill for AWS DevOps Agent. Given a domain ARN (or domain name + region), it runs control-plane API checks across cluster health, shard strategy, performance, security posture, and cost optimization, then produces a structured findings report with prioritized recommendations. + +## What it does + +Analyzes an OpenSearch domain across five categories: + +1. **Cluster Health & Configuration** — status, node/AZ balance, dedicated masters, upgrade eligibility +2. **Storage & Shard Strategy** — EBS volume type, shard density, active vs. total shards, free storage +3. **Performance** — JVM pressure, CPU, search/indexing latency, HTTP errors, cluster status history +4. **Security & Access** — encryption at rest, node-to-node encryption, HTTPS enforcement, network exposure, access policy, fine-grained access control +5. **Cost Optimization** — instance right-sizing, storage tiering, reserved instance coverage + +It is **100% control-plane / API-driven** — no data-plane access (no `_cluster/*` calls, no index reads), which makes it compatible with DevOps Agent without a custom MCP server. + +## Prerequisites + +**IAM permissions** the DevOps Agent execution role needs (all read-only): + +- `es:DescribeDomain` +- `es:DescribeDomainHealth` +- `es:DescribeDomainConfig` +- `es:GetCompatibleVersions` +- `es:DescribeReservedInstances` +- `es:ListDomainNames` +- `es:ListTags` +- `cloudwatch:GetMetricData` + +Most are covered by `AIDevOpsAgentAccessPolicy`; attach the supplemental policy for any gaps. + +**AWS resources:** one or more OpenSearch Service domains in the target account/region. + +## How to use it with DevOps Agent + +1. Zip this skill directory and upload it to your DevOps Agent Space. +2. Select relevant subagents (Chat, Investigations / Incident RCA). +3. Prompt in natural language without naming the skill, e.g.: + - "Is my OpenSearch domain healthy? ARN: `arn:aws:es:us-east-1:123456789012:domain/my-domain`" + - "Review the cluster configuration and security posture of this OpenSearch domain." + - "Are there any cost savings available on my OpenSearch setup?" +4. Review the agent's reasoning trace to confirm the skill activated and the checks ran. + +If the agent does not invoke the skill, refine the `description` field in `SKILL.md` (see "Optimizing description" in the Agent Skills specification). + +## Non-production disclaimer + +> ⚠️ This skill is sample code, not intended for production use without additional review and testing. Users should validate in a non-production environment first. + +## Maintainers + +- genealpe +- prasadnu diff --git a/skills/analytics-opensearch-expertise/SKILL.md b/skills/analytics-opensearch-expertise/SKILL.md new file mode 100644 index 0000000..e4a6bf7 --- /dev/null +++ b/skills/analytics-opensearch-expertise/SKILL.md @@ -0,0 +1,395 @@ +--- +name: analytics-opensearch-expertise +description: "Amazon OpenSearch Service domain health assessment. Performs read-only, API-driven checks against a customer's OpenSearch domain(s) covering cluster health, node/shard configuration, performance metrics, security posture, and cost optimization signals. Activate this skill for requests about OpenSearch or Elasticsearch domain health, cluster review, domain assessment, performance, security posture, or cost optimization. Given a domain ARN (or name + region), it produces a structured findings report with prioritized recommendations. All checks use read-only AWS control-plane APIs (es:Describe*, es:GetCompatibleVersions, es:DescribeReservedInstances, cloudwatch:GetMetricData) — no data-plane access required." +metadata: + version: "2.6" + author: genealpe +--- + +# OpenSearch Domain Health Assessment + +## Overview + +This skill performs a comprehensive, read-only health assessment of Amazon OpenSearch Service domains using AWS APIs available through the customer's local AWS profile. It produces a structured report with findings and actionable recommendations. + +> **Scope:** CLI-agent compatible. No external packages, no IDE workspace, no human review gates. + +## Trigger Keywords + +`opensearch`, `opensearch health`, `domain health`, `cluster review`, `opensearch review`, `opensearch assessment`, `opensearch check`, `open search`, `opensearch performance`, `opensearch security`, `opensearch cost` + +## Prerequisites + +- AWS CLI profile configured with read-only access to the target account +- IAM permissions required: + - `es:DescribeDomain` + - `es:DescribeDomainHealth` + - `es:GetCompatibleVersions` + - `es:DescribeReservedInstances` + - `es:ListDomainNames` + - `cloudwatch:GetMetricData` + +## Execution Flow + +### Input + +**Required:** Domain ARN (e.g., `arn:aws:es:us-west-2:123456789012:domain/my-domain`) + +Parse the ARN to extract: +- **Region:** field 4 (e.g., `us-west-2`) +- **Account ID:** field 5 (e.g., `123456789012`) +- **Domain name:** after `domain/` in field 6 + +**Fallback / Target Resolution (apply these rules exactly — never silently substitute a target):** +1. **Exact match:** the ARN (or name + region) resolves via `DescribeDomain` → proceed with the assessment. +2. **No domain specified:** run `ListDomainNames` in the user's likely region(s) (cross-region scan if needed) and **present the discovered domains as options for the user to choose**. Do not pick one yourself unless exactly one domain exists in the entire account — state that you did so. +3. **Specified but NOT found:** search for candidates (`ListDomainNames` in the ARN's region, then other regions). Then **STOP and confirm — do NOT assess**: + - Report clearly: "Domain {given} not found in {region}." + - List the closest candidate(s) with their regions (e.g., name differs by a suffix, or same name in another region). + - Ask the user to confirm which domain to assess. **Never run the assessment against a domain the user did not name without their explicit confirmation** — in accounts with similarly-named sibling domains (e.g., `prod-es-content` / `prod-es-users`), a silently substituted target produces a confident report about the wrong domain, which is worse than an error. +4. If the user provides only a domain name with no region, ask for the region (or apply rule 3's candidate search). + +### Steps + +``` +1. Parse domain ARN → region, account, domain name +2. Resolve the target (Target Resolution rules above) — confirm with the user if not an exact match +3. Describe domain configuration (DescribeDomain) +4. Describe domain health (DescribeDomainHealth) +5. Pull CloudWatch metrics (last 24h) +6. Check upgrade eligibility (GetCompatibleVersions; ListVersions fallback) +7. Check reserved instance coverage (DescribeReservedInstances) +8. Analyze findings against thresholds +9. Load the Remediation Reference — read_skill_resource(skill_id='analytics-opensearch-expertise', + path='references/remediation-reference.md') — MANDATORY before writing recommendations +10. Generate report with recommendations +``` + +## API Compatibility (DevOps Agent tooling) + +The DevOps Agent AWS tool exposes the LEGACY Elasticsearch-compatible client. Call legacy +operation names FIRST; fall back to the modern names only if the legacy call fails: + +| Purpose | Try first (legacy) | Modern equivalent | +|---|---|---| +| Domain config | `describe_elasticsearch_domain` | `describe_domain` | +| Version catalog | `list_elasticsearch_versions` | `list_versions` | +| Upgrade targets | `get_compatible_elasticsearch_versions` | `get_compatible_versions` | +| Reserved instances | `describe_reserved_elasticsearch_instances` | `describe_reserved_instances` | + +- IAM note: the execution role's managed policy grants `es:Describe*` and `es:List*` — + Describe/List operations succeed under either naming; `Get*` actions (BOTH GetCompatible + variants) are typically denied. This is exactly why the 1.4 ListVersions fallback exists. +- `DescribeDomainHealth` has NO legacy equivalent. If it fails, do NOT skip checks 2.2/2.3 — + use CloudWatch instead: `Shards.active` (Maximum) = active shards, `Shards.unassigned` + (Maximum) = unassigned, TotalShards ≈ active + unassigned; `ClusterStatus.green/yellow/red` + for 1.1/3.7. Never report the checks as SKIPPED when this fallback is available. +- All CloudWatch metrics in this skill: namespace `AWS/ES`, dimensions `DomainName` and + `ClientId` (the account ID). + +## Checks & Decision Logic + +### Category 1: Cluster Health & Configuration + +#### 1.1 Cluster Status +- **API:** `DescribeDomainHealth` +- **Field:** `ClusterHealth` (green | yellow | red) +- **Logic:** + - GREEN → ✅ All primary and replica shards assigned + - YELLOW → ⚠️ All primaries assigned, some replicas unassigned + - RED → 🔴 Some primary shards unassigned — data loss risk + +#### 1.2 Node Count & AZ Balance +- **API:** `DescribeDomain` → `ClusterConfig` +- **Fields:** `InstanceCount`, `ZoneAwarenessEnabled`, `AvailabilityZoneCount` +- **Logic:** + - If `ZoneAwarenessEnabled == true`: + - `InstanceCount % AvailabilityZoneCount != 0` → ⚠️ "Data node count ({count}) is not a multiple of AZ count ({az_count}). This causes uneven shard distribution across AZs, reducing fault tolerance." + - Recommendation: "Adjust data node count to a multiple of {az_count} (e.g., {nearest_multiple})." + - If `ZoneAwarenessEnabled == false` and `InstanceCount > 1`: + - ⚠️ "Multi-AZ is not enabled despite having {count} data nodes. Single-AZ deployment risks full cluster unavailability during AZ failure." + +#### 1.3 Dedicated Master Configuration +- **API:** `DescribeDomain` → `ClusterConfig` +- **Fields:** `DedicatedMasterEnabled`, `DedicatedMasterType`, `DedicatedMasterCount` +- **Logic:** + - `DedicatedMasterEnabled == false` AND `InstanceCount >= 3` → ⚠️ "No dedicated master nodes. Clusters with 3+ data nodes should use dedicated masters for stability." + - `DedicatedMasterCount == 2` → ⚠️ "Even number of master nodes ({count}) risks split-brain. Use 3 dedicated master nodes." + - `DedicatedMasterCount < 3` → ⚠️ "Fewer than 3 master nodes reduces fault tolerance." + +#### 1.4 Engine Version & Upgrade Eligibility +- **API:** `GetCompatibleVersions` with `DomainName` +- **Fallback (use when `GetCompatibleVersions` returns AccessDeniedException — do NOT skip):** + call `ListVersions` and derive upgrade candidates using these two rules: + 1. Direct upgrade candidates = versions in `ListVersions` with the SAME major version as the + domain and a LATER minor version (e.g., a 2.11 domain → every later 2.x in the region). + 2. Next-major path = via the LATEST available minor of the current major (compute it from + `ListVersions`; never hard-code a specific version number). + Label fallback results as "derived upgrade candidates (ListVersions + documented upgrade-path + rules)" and cite https://docs.aws.amazon.com/opensearch-service/latest/developerguide/version-migration.html + for the reader. Only mark 1.4 SKIPPED if BOTH APIs are denied. + **Note placement:** the fallback provenance note belongs INSIDE the 1.4 finding ONLY — do NOT + surface it in the Executive Summary or any report-level note box, and do NOT name API errors + (e.g., AccessDeniedException) anywhere in the report. Phrase it customer-friendly: "Upgrade + candidates were derived from the available version listings and are not API-confirmed — verify + eligible target versions in the console before planning an upgrade." + Render candidates as a SET of direct single-hop targets (comma-separated) — NEVER as an arrow chain (NOT "2.13 → 2.15 → 2.19"): every listed same-major version is directly reachable in ONE hop. Any upgrade recommendation MUST propose a single hop to the latest same-major version (plus one more hop for a major jump). Do NOT invent sequential multi-hop plans or per-hop time estimates. +- **Logic:** + - If compatible target versions exist → ℹ️ "Current version: {current}. Upgrades available: {targets}." + - If current version is more than 2 releases behind the latest same-major version → ⚠️ "Consider upgrading for performance improvements and security patches." + - **Counting rule (apply exactly):** "releases behind" = the COUNT of later same-major versions returned by `ListVersions`/`GetCompatibleVersions` (e.g., 2.11 with {2.13, 2.15, 2.17, 2.19} available = 4 releases behind) — NEVER the numeric delta between version numbers (2.19 − 2.11 is NOT "8 versions behind"). + +### Category 2: Storage & Shard Strategy + +#### 2.1 EBS Configuration +- **API:** `DescribeDomain` → `EBSOptions` +- **Fields:** `VolumeType`, `VolumeSize`, `Iops`, `Throughput` +- **Logic:** + - `VolumeType == "gp2"` → ⚠️ "Using gp2 volumes. gp3 offers better price/performance with configurable IOPS and throughput." + - Report total cluster storage: `VolumeSize × InstanceCount` + +#### 2.2 Shard Count & Density +- **API:** `DescribeDomainHealth` +- **Fields:** `TotalShards`, `ActiveShards` +- **Compute:** `shards_per_node = TotalShards / InstanceCount` +- **Version-dependent limits (from `DescribeDomain` → `EngineVersion`):** + - OpenSearch ≤ 2.15: Hard limit = **1,000 shards/node** + - OpenSearch 2.17+: Hard limit = **1,000 shards per 16 GB of data node heap** +- **Heap estimation (MUST use the real instance type, not an assumed value):** estimate per-node JVM heap from the data node instance type: `heap_GiB ≈ min(32, instance_RAM_GiB / 2)` (e.g., t3.medium.search 4 GiB RAM → ~2 GiB heap; r6g.xlarge.search 32 GiB RAM → ~16 GiB heap). + **Instance RAM reference (GiB):** t3.small=2 · t3.medium=4 · c6g.large=4 · m6g.large=8 · + c6g.xlarge=8 · r6g.large=16 · m6g.xlarge=16 · c6g.2xlarge=16 · r6g.xlarge=32 · + m6g.2xlarge=32 · r6g.2xlarge=64 (heap caps at 32 GiB). For unlisted types use the EC2 + equivalent's memory — never guess silently; name the assumed RAM in the report. +- **Best-practice threshold:** `bp_limit = 25 × heap_GiB` shards/node. +- **Logic (apply these exact thresholds — do not substitute your own; the verdict tier is determined ONLY by which line matches, never by judgment):** + - `shards_per_node > hard_limit` → 🔴 "Exceeds hard shard limit for engine version {version}." + - `shards_per_node > 4 × bp_limit` → 🔴 "Shard density ({shards_per_node}/node) is more than 4× the best-practice limit ({bp_limit}/node for ~{heap_GiB} GiB heap). Severe heap/GC and recovery risk." (This tier is 🔴, not ⚠️ — do not downgrade.) + - `shards_per_node > bp_limit` → ⚠️ "Shard density ({shards_per_node}/node) exceeds best-practice threshold of 25 shards/GiB heap ({bp_limit}/node for ~{heap_GiB} GiB heap)." + - `shards_per_node > 0.5 × bp_limit` → ℹ️ "Moderate shard density. Monitor JVM heap pressure." + - `shards_per_node <= 0.5 × bp_limit` → ✅ PASS. (Verbatim: at or below half the best-practice limit the verdict is PASS, never ℹ️ — low density is not a finding.) + +#### 2.3 Active Shards vs. Total Shards +- **API:** `DescribeDomainHealth` +- **Fields:** `ActiveShards`, `TotalShards` +- **Expected:** `ActiveShards == TotalShards` (all shards placed and serving) +- **Logic:** + - `ActiveShards < TotalShards`: + - Delta = `TotalShards - ActiveShards` + - ⚠️ "{delta} shards are not active. Non-active shards may be:" + - "• **Initializing** — being allocated after a node restart or scaling event (transient, typically resolves in minutes)" + - "• **Relocating** — being moved during a rebalance operation (transient)" + - "• **Unassigned** — cannot be placed on any node (persistent problem). Common causes: insufficient disk space, allocation filtering rules, AZ awareness constraints with unbalanced node counts, or max shards per node limit reached." + - Recommendation: "Check `_cluster/allocation/explain` for unassigned shard root cause." + +#### 2.4 Free Storage Space +- **API:** CloudWatch `FreeStorageSpace` (Minimum statistic, last 24h) +- **CRITICAL — the Minimum statistic is PER-NODE (the lowest single node's free MB), NOT a + cluster-wide total. Compute free % against ONE node's volume, never the cluster total:** + `free_pct = FreeStorageSpace_min_MB / (VolumeSize_GiB × 1024) × 100` + (Dividing the per-node Minimum by `VolumeSize × InstanceCount` understates free space by a + factor of InstanceCount — this is wrong. Cross-check: the Sum statistic ÷ (VolumeSize × + InstanceCount × 1024) should give approximately the same percentage.) +- **Note:** usable space per node is less than the raw volume size due to OS/service reserved + overhead — an empty node typically shows ~72% free, not ~97%. +- **Logic (apply these exact thresholds verbatim):** + - `free_pct < 10%` → 🔴 "Free storage below 10%. High watermark may trigger read-only mode." + - `free_pct < 20%` → ⚠️ "Free storage below 20%. Approaching low watermark. Plan capacity increase." + - `free_pct < 30%` → ℹ️ "Storage utilization above 70%. Monitor trend." + - `free_pct >= 30%` → ✅ PASS. + - If the metric returns NO datapoints (new domain, or configuration change in progress) → + SKIPPED "FreeStorageSpace metric unavailable" — never silently omit the check. + +### Category 3: Performance Metrics + +> **Note on percentiles:** CloudWatch metrics for OpenSearch (`SearchLatency`, `IndexingLatency`) report aggregate values, not per-request percentiles. We use the **Maximum** statistic over 1-minute periods as a "worst-case" proxy. True p50/p99 latency requires OpenSearch slow-log analysis or UltraWarm query insights, which is outside the scope of API-only checks. + +**Time range:** Last 24 hours, 1-minute period granularity. + +#### 3.1 JVM Memory Pressure +- **Metric:** `JVMMemoryPressure`, Statistic: `Maximum` — **the verdict MUST be computed from the + Maximum statistic. Never substitute Average (an idle cluster commonly shows Avg ~43% while Max + sustains ~74% — these produce different verdicts). Report Average only as context.** +- **Logic:** + - Max > 92% → 🔴 "JVM heap pressure critically high ({value}%). Cluster is at risk of circuit breaker trips and OOM. Consider scaling up instance type or reducing shard count." + - Max > 80% → ⚠️ "JVM heap pressure elevated ({value}%). Approaching GC thrashing threshold." + - Max > 70% → ℹ️ "JVM heap pressure moderate. Monitor trend." + +#### 3.2 CPU Utilization +- **Metric:** `CPUUtilization`, Statistic: `Average` (sustained) and `Maximum` (spikes) +- **Logic:** + - Avg > 80% sustained (>50% of datapoints) → ⚠️ "Sustained high CPU. Cluster may be undersized for current workload." + - Max > 95% → ℹ️ "CPU spikes observed. Check for expensive queries or bulk indexing bursts." + +#### 3.3 Search Latency +- **Metric:** `SearchLatency`, Statistic: `Maximum` (worst-case proxy) +- **Logic:** + - Max > 500ms → ℹ️ "Worst-case search latency {value}ms — elevated spikes observed. Investigate slow queries if user-facing impact is reported." + - Report average as baseline context. + - **Informational only — search latency NEVER renders ⚠️ or 🔴.** The 500ms marker is a heuristic: AWS does not define a canonical latency threshold; acceptable latency is workload-dependent, so frame findings against the customer's own baseline. + +#### 3.4 Indexing Latency +- **Metric:** `IndexingLatency`, Statistic: `Maximum` (worst-case proxy) +- **Logic:** + - Max > 500ms → ℹ️ "Worst-case indexing latency {value}ms. Check merge pressure or refresh interval if ingestion throughput is impacted." + - Report average as baseline context. + - **Informational only — indexing latency NEVER renders ⚠️ or 🔴.** Heuristic marker; frame against the workload's own baseline. + +#### 3.5 Search & Indexing Rate +- **Metrics:** `SearchRate` (Sum), `IndexingRate` (Sum) +- **Purpose:** Workload characterization — reported for context, no threshold alerts. +- **Output:** "Search rate: ~{avg_per_min} req/min. Indexing rate: ~{avg_per_min} docs/min." + +#### 3.6 HTTP Errors +- **Metrics:** `4xx` (Sum), `5xx` (Sum) +- **Logic (the tier is determined ONLY by 5xx; 4xx volume never raises the tier):** + - `5xx > 0` → ⚠️ "{count} server errors (5xx) in last 24h. Indicates cluster-side failures (overload, circuit breaker, shard failures)." + - `5xx == 0` AND `4xx > 0` → ℹ️ "{count} client errors (4xx). May indicate malformed queries, auth issues, or unauthenticated probes against a public endpoint." (Verbatim: with zero 5xx this is ℹ️ regardless of 4xx volume — never ⚠️. Note spikes/patterns as context only.) + - `5xx == 0` AND `4xx == 0` → ✅ PASS. + +#### 3.7 Cluster Status History +- **Metrics:** `ClusterStatus.red` (Maximum), `ClusterStatus.yellow` (Maximum) +- **Logic:** + - Any `red == 1` datapoints → 🔴 "Cluster entered RED status {count} times in last 24h." + - Any `yellow == 1` datapoints (and currently not yellow) → ℹ️ "Cluster experienced YELLOW status {count} times in last 24h (now recovered)." + - If the cluster is CURRENTLY yellow (1.1 is already ⚠️), still report the 24h yellow + datapoint count here as history context — never omit 3.7. + +### Category 4: Security & Access + +> **SEVERITY OVERRIDE PROHIBITION (apply exactly):** The verdict tier for every check in this +> document comes ONLY from the Logic line that matches the observed values — NEVER from general +> security judgment. Several security checks below are deliberately rated ⚠️ (not 🔴) by owner +> decision: a single disabled control in isolation is a Warning; only the fully-compounded 4.4 +> case is Critical. Rendering 4.1, 4.2, 4.3, 4.5, or 4.6 as 🔴/CRITICAL is a SPEC VIOLATION even +> if it feels more correct — if your judgment disagrees with a Logic line, the Logic line wins. +> You may add ONE sentence of context inside the finding; you may never change the verdict tier, +> the emoji, or the severity word. This applies to the matrix, the executive summary, section +> headings, and prose equally. + +#### 4.1 Encryption at Rest +- **API:** `DescribeDomain` → `EncryptionAtRestOptions.Enabled` +- **Logic:** + - `false` → ⚠️ "Encryption at rest is disabled. Data on disk is unencrypted — non-compliant with most regulatory frameworks." (⚠️ by owner decision — do NOT escalate to 🔴.) + +#### 4.2 Node-to-Node Encryption +- **API:** `DescribeDomain` → `NodeToNodeEncryptionOptions.Enabled` +- **Logic:** + - `false` → ⚠️ "Node-to-node encryption disabled. Inter-node traffic is unencrypted." (⚠️ by owner decision — do NOT escalate to 🔴.) + +#### 4.3 HTTPS Enforcement +- **API:** `DescribeDomain` → `DomainEndpointOptions.EnforceHTTPS` +- **Logic:** + - `false` → ⚠️ "HTTPS not enforced. Clients can connect over HTTP (plaintext)." + - Check `TLSSecurityPolicy` — if not `Policy-Min-TLS-1-2-2019-07` or newer → ℹ️ "TLS policy allows older protocol versions." + +#### 4.4 Network Exposure +- **API:** `DescribeDomain` → `VPCOptions` +- **Logic (deterministic compound-risk escalation — apply exactly):** + - If `VPCOptions` is empty/null AND `EnforceHTTPS == false` AND `AdvancedSecurityOptions.Enabled == false` → 🔴 "Domain uses a public endpoint with BOTH transport and access controls disabled — fully compounded exposure. Migrate to VPC or close both gaps." (This is an AND — BOTH controls must be disabled. A single weak control does NOT escalate 4.4; it is already flagged ⚠️ by its own check: HTTPS by 4.3, FGAC by 4.6. The fully-compounded case is the ONLY security condition rated 🔴.) + - If `VPCOptions` is empty/null (any other combination) → ⚠️ "Domain uses a public endpoint{; weak control: {name the single false one, if any}}. Consider VPC deployment for network isolation." + - If VPC deployed → ✅ Report VPC ID, subnet IDs, security group IDs. + +#### 4.5 Access Policy Analysis +- **API:** `DescribeDomain` → `AccessPolicies` (JSON string) +- **Logic:** Parse the resource-based policy: + - Principal == `"*"` with no IP condition → ⚠️ "Access policy allows unauthenticated access from any IP." + - Principal == `"*"` with IP condition → ℹ️ "Access restricted by IP/CIDR but no IAM authentication." + - Specific AWS principals → ✅ Report principal list. + +#### 4.6 Fine-Grained Access Control (FGAC) +- **API:** `DescribeDomain` → `AdvancedSecurityOptions` +- **Fields:** `Enabled`, `InternalUserDatabaseEnabled` +- **Logic:** + - `Enabled == false` → ⚠️ "Fine-grained access control disabled. No index-level or document-level permissions." (⚠️ by owner decision — do NOT escalate to 🔴.) + - `InternalUserDatabaseEnabled == true` → ℹ️ "Internal user database enabled. Consider SAML or IAM-based auth for production." + +### Category 5: Cost Optimization Signals + +#### 5.1 Instance Right-Sizing +- **Inputs:** Instance type (from config) + CPU avg + JVM pressure +- **Logic (BOTH conditions MUST hold — this is a conjunctive rule):** + - CPU avg < 20% AND JVM pressure max < 50% → ℹ️ "Low utilization suggests potential over-provisioning. Consider downsizing instance type." + - If JVM pressure ≥ 50%, do NOT emit this finding regardless of CPU — the heap is doing real work even if CPU is idle. + - Both high → already flagged in performance section. + +#### 5.2 Storage Tiering Opportunity +- **API:** `DescribeDomain` → `ClusterConfig.WarmEnabled`, `ColdStorageOptions.Enabled` +- **Logic:** + - `WarmEnabled == false` → ℹ️ "UltraWarm not enabled. If you have infrequently accessed indices (e.g., >30 days old), UltraWarm can reduce cost by ~80% for those indices." + - `ColdStorageEnabled == false` AND `WarmEnabled == true` → ℹ️ "Cold storage not enabled. For rarely accessed data, cold storage offers lowest-cost option." + +#### 5.3 Reserved Instance Coverage +- **API:** `DescribeReservedInstances` +- **Logic:** + - No active RIs → ℹ️ "No reserved instances. For steady-state workloads, RIs can save 30-50% vs. on-demand." (An EMPTY result is NOT an error — empty means no RIs.) + - If the API returns AccessDenied → SKIPPED with the reason stated. + - Compare RI instance type/count vs. current domain config → report coverage gap. + +### Remediation Reference rule (apply exactly) +For EVERY finding rated 🔴, ⚠️, or ℹ️, the Prioritized Recommendations section MUST include that check's entry from `references/remediation-reference.md` (loaded in Execution Flow step 9 via read_skill_resource): copy the "Why it matters" line and "Resolve" steps verbatim (you may instantiate observed values, e.g. the actual master count), and include the "Dive deeper" link(s) EXACTLY as written. NEVER emit a documentation link that is not present in the Remediation Reference — do not construct, recall, or infer URLs from any other source. PASS/✅ checks get no remediation entry. + +## Output Format + +> **MANDATORY COVERAGE RULE:** The report MUST evaluate and account for EVERY check in this document (1.1–1.4, 2.1–2.4, 3.1–3.7, 4.1–4.6, 5.1–5.3 — 24 checks total). No check may be silently omitted. +> **ID FIDELITY:** matrix rows MUST use these exact IDs with these exact meanings — never renumber, +> split, merge, or invent checks: 1.1 Cluster Status · 1.2 Node Count & AZ Balance · 1.3 Dedicated +> Masters · 1.4 Engine Version · 2.1 EBS Configuration · 2.2 Shard Count & Density · 2.3 Active vs +> Total Shards · 2.4 Free Storage · 3.1 JVM Memory Pressure · 3.2 CPU Utilization · 3.3 Search +> Latency · 3.4 Indexing Latency · 3.5 Search & Indexing Rate · 3.6 HTTP Errors · 3.7 Cluster +> Status History · 4.1 Encryption at Rest · 4.2 Node-to-Node Encryption · 4.3 HTTPS Enforcement · +> 4.4 Network Exposure · 4.5 Access Policy · 4.6 FGAC · 5.1 Instance Right-Sizing · 5.2 Storage +> Tiering · 5.3 Reserved Instances. (Splitting 1.1 into red/yellow/green rows, or using "1.4" for +> service-software patches, are documented failure modes — do not reproduce them.) +> **CATEGORY ROLL-UP:** each category's executive-summary status = the WORST finding in it: +> any 🔴 → Critical; else any ⚠️ → Warning; else all ✅/ℹ️ → Healthy. Never judgment-based. Apply the exact thresholds written in each check's Logic section verbatim — do not substitute your own thresholds or severity levels, and do not round a verdict up or down a tier based on judgment: the tier is determined solely by which Logic line matches. If a check cannot be evaluated (API error, no data), it MUST appear in the Check Coverage Matrix as SKIPPED with the reason. For every CloudWatch-based check (2.4, 3.1–3.7, 5.1), the matrix's Observed Value column MUST name the statistic used (e.g., "Max=74.3% (Maximum stat)") — it must match the statistic named in that check's spec. Before finishing, count the matrix rows: if the count is not exactly 24, the report is incomplete — fix it before responding. + +Structure the report as: + +```markdown +# OpenSearch Domain Health Assessment +## Domain: {domain_name} +**Region:** {region} | **Engine:** {engine_version} | **Assessed:** {timestamp} + +## Summary +| Category | Status | Findings | +|----------|--------|----------| +| Cluster Health | 🟢/🟡/🔴 | {one-line summary} | +| Storage & Shards | 🟢/🟡/🔴 | {one-line summary} | +| Performance | 🟢/🟡/🔴 | {one-line summary} | +| Security | 🟢/🟡/🔴 | {one-line summary} | +| Cost Optimization | 🟢/🟡/⚪ | {one-line summary} | + +## Detailed Findings + +### 🔴 Critical (act now) +{findings} + +### ⚠️ Warnings (plan action) +{findings} + +### ℹ️ Informational +{findings} + +### ✅ Passing Checks +{list of checks that passed} + +## Recommendations (prioritized) +1. {highest priority} +2. ... + +## Raw Data Reference +{key metrics and config values for verification} + +## Check Coverage Matrix (REQUIRED — one row per check, all 24, in order) +| Check | Verdict | Observed value | Threshold applied | +|-------|---------|----------------|-------------------| +| 1.1 Cluster status | 🟢/⚠️/🔴 | {value} | {rule} | +| 1.2 Node/AZ balance | ... | ... | ... | +| ... (every check through 5.3 — SKIPPED rows must state the reason) | +``` + +## Error Handling + +- If `DescribeDomainHealth` returns `Processing` → domain is being modified. Report current state with caveat. +- If any API returns `AccessDeniedException` → report which check was skipped and what permission is needed. +- If CloudWatch returns no datapoints → report "No data available for {metric}. Domain may be newly created or idle." diff --git a/skills/analytics-opensearch-expertise/evals/benchmark.json b/skills/analytics-opensearch-expertise/evals/benchmark.json new file mode 100644 index 0000000..59f7d0d --- /dev/null +++ b/skills/analytics-opensearch-expertise/evals/benchmark.json @@ -0,0 +1,828 @@ +{ + "skill_name": "analytics-opensearch-expertise", + "skill_path": "/home/genealpe/skills/analytics-opensearch-expertise", + "eval_count": 7, + "runs_per_eval": 1, + "metadata": { + "timestamp": "2026-08-12T11:58:36Z" + }, + "runs": [ + { + "eval_id": "opensearch-health-categories", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-health-categories", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Cluster Health' or contains 'cluster health'", + "passed": true, + "evidence": "OR satisfied by: contains 'cluster health' \u2014 Substring found: 'cluster health'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Storage' or contains 'Shard'", + "passed": true, + "evidence": "OR satisfied by: contains 'storage' \u2014 Substring found: 'storage'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Performance' or contains 'performance'", + "passed": true, + "evidence": "OR satisfied by: contains 'performance' \u2014 Substring found: 'performance'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": true, + "evidence": "OR satisfied by: contains 'security' \u2014 Substring found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": true, + "evidence": "OR satisfied by: contains 'cost' \u2014 Substring found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 15.636458222026704 + }, + "raw_output": "> The five assessment categories are:\n\n1. Cluster Health & Configuration \u2014 cluster status, node count/AZ balance, dedicated masters, engine version and upgrade eligibility\n2. Storage & Shard Strategy \u2014 EBS configuration, shard count/density, active vs. total shards, free storage space\n3. Performance Metrics \u2014 JVM memory pressure, CPU, search/indexing latency and rates, HTTP errors, cluster status history\n4. Security & Access \u2014 encryption at rest, node-to-node encryption, HTTPS enforcement, network exposure, access policy, fine-grained access control\n5. Cost Optimization Signals \u2014 instance right-sizing, storage tiering (UltraWarm/cold), reserved instance coverage" + }, + "without_skill": { + "eval_id": "opensearch-health-categories", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Cluster Health' or contains 'cluster health'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'cluster health': Substring not found: 'cluster health'; contains 'cluster health': Substring not found: 'cluster health'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Storage' or contains 'Shard'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'storage': Substring not found: 'storage'; contains 'shard': Substring not found: 'shard'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Performance' or contains 'performance'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'performance': Substring not found: 'performance'; contains 'performance': Substring not found: 'performance'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Security' or contains 'security'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'security': Substring not found: 'security'; contains 'security': Substring not found: 'security'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Cost' or contains 'cost'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'cost': Substring not found: 'cost'; contains 'cost': Substring not found: 'cost'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 16.34111326900893 + }, + "raw_output": "> I don't have a skill that assesses OpenSearch domains. The skills available to me right now are:\n\n- **2x2 Note Collector** \u2014 captures/categorizes notes for your bi-weekly 2x2 report\n- **DLS Program Manager** \u2014 Analytics TFC Data Learning Series program management\n- **artifactory-design** \u2014 HTML artifact design guidance\n\nNone of them define assessment categories for an OpenSearch domain, so I can't quote the five categories without inventing them. If you have an OpenSearch review/assessment skill (e.g., a SKILL.md in ~/.kiro/skills/), point me to its path or drop it in, and I'll list the categories exactly as written." + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "opensearch-required-apis", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-required-apis", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'DescribeDomain'", + "passed": true, + "evidence": "Substring found: 'describedomain'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'DescribeDomainHealth'", + "passed": true, + "evidence": "Substring found: 'describedomainhealth'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'GetCompatibleVersions'", + "passed": true, + "evidence": "Substring found: 'getcompatibleversions'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'DescribeReservedInstances'", + "passed": true, + "evidence": "Substring found: 'describereservedinstances'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'ListDomainNames'", + "passed": true, + "evidence": "Substring found: 'listdomainnames'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'GetMetricData'", + "passed": true, + "evidence": "Substring found: 'getmetricdata'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 18.775910616008332 + }, + "raw_output": "> From the skill's prerequisites, it needs these read-only IAM permissions:\n\n- es:DescribeDomain \u2014 domain configuration (cluster config, EBS, security options)\n- es:DescribeDomainHealth \u2014 cluster status, shard counts\n- es:GetCompatibleVersions \u2014 upgrade eligibility\n- es:DescribeReservedInstances \u2014 RI coverage\n- es:ListDomainNames \u2014 target resolution / domain discovery\n- cloudwatch:GetMetricData \u2014 performance and storage metrics (namespace AWS/ES)\n\nA few practical notes from the skill:\n\n- Legacy API names: in DevOps Agent tooling the skill tries the legacy Elasticsearch-compatible operations first \u2014 DescribeElasticsearchDomain, ListElasticsearchVersions, GetCompatibleElasticsearchVersions, DescribeReservedElasticsearchInstances \u2014 falling back to the modern names. A broad es:Describe* + es:List* grant covers both namings.\n- Get* actions may be denied under a Describe/List-only policy \u2014 the skill has a built-in fallback for that: it derives upgrade candidates from es:ListVersions instead of GetCompatibleVersions, so 1.4 still runs.\n- Everything is control-plane only \u2014 no data-plane (cluster endpoint) access is required." + }, + "without_skill": { + "eval_id": "opensearch-required-apis", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'DescribeDomain'", + "passed": false, + "evidence": "Substring not found: 'describedomain'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'DescribeDomainHealth'", + "passed": false, + "evidence": "Substring not found: 'describedomainhealth'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'GetCompatibleVersions'", + "passed": false, + "evidence": "Substring not found: 'getcompatibleversions'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'DescribeReservedInstances'", + "passed": false, + "evidence": "Substring not found: 'describereservedinstances'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'ListDomainNames'", + "passed": false, + "evidence": "Substring not found: 'listdomainnames'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'GetMetricData'", + "passed": false, + "evidence": "Substring not found: 'getmetricdata'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 16.426152881002054 + }, + "raw_output": "> None. The 2x2 Note Collector skill doesn't touch AWS at all \u2014 it needs no AWS APIs, credentials, or IAM permissions.\n\nEverything it does runs through:\n- Local filesystem \u2014 reading/writing JSON and markdown under /local/home/genealpe/obsidian-vault/2x2_notes/\n- Outlook MCP tools \u2014 calendar_view, email_folders, email_search (authenticated via your Microsoft 365 identity, not IAM)\n- Slack MCP tools \u2014 search, get_messages (Slack token, not IAM)\n\nSo there's no IAM policy to scope or read-only role to configure for it." + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "opensearch-arn-parsing", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-arn-parsing", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'region'", + "passed": true, + "evidence": "Substring found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'account'", + "passed": true, + "evidence": "Substring found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'domain'", + "passed": true, + "evidence": "Substring found: 'domain'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /field|segment|position|part|colon/", + "passed": true, + "evidence": "Regex /field|segment|position|part|colon/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 21.15945057402132 + }, + "raw_output": "> An OpenSearch domain ARN is a colon-delimited string with six fields:\n\narn:aws:es:us-west-2:123456789012:domain/my-domain\n 1 2 3 4 5 6\n\n\nSplitting on ::\n\n- **Field 4 \u2192 Region** \u2014 e.g., us-west-2\n- **Field 5 \u2192 Account ID** \u2014 e.g., 123456789012\n- **Field 6 \u2192 Domain name** \u2014 the last field is domain/my-domain; the skill takes everything after the domain/ prefix, yielding my-domain\n\nEquivalent shell one-liner:\n\nbash\nARN=\"arn:aws:es:us-west-2:123456789012:domain/my-domain\"\nREGION=$(echo \"$ARN\" | cut -d: -f4)\nACCOUNT=$(echo \"$ARN\" | cut -d: -f5)\nDOMAIN=$(echo \"$ARN\" | cut -d: -f6 | sed 's|^domain/||')\n\n\nTwo things worth noting:\n\n- The service field (field 3) is es, not opensearch \u2014 Amazon OpenSearch Service kept the legacy Elasticsearch ARN namespace, which is also why the skill tries legacy API names like describe_elasticsearch_domain first.\n- After parsing, the skill doesn't trust the ARN blindly: it runs DescribeDomain to confirm the domain actually exists in that region. If it doesn't resolve, it searches for candidates via ListDomainNames (same region, then others) and stops to ask you which domain to assess rather than substituting a similar-looking one." + }, + "without_skill": { + "eval_id": "opensearch-arn-parsing", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'region'", + "passed": true, + "evidence": "Substring found: 'region'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'account'", + "passed": true, + "evidence": "Substring found: 'account'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'domain'", + "passed": true, + "evidence": "Substring found: 'domain'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "matches regex /field|segment|position|part|colon/", + "passed": true, + "evidence": "Regex /field|segment|position|part|colon/ matched", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 37.29119509001612 + }, + "raw_output": "Reading directory: /home/genealpe/.kiro/skills (using tool: read, max depth: 1, max entries: 1000, excluding: defaults)> I couldn't verify which skill you mean \u2014 none of the skills loaded in my context (2x2-note-collector, dls-program-manager, artifactory-design) take a domain ARN, and I was blocked from listing your skills directory to find it. So here's how domain ARN parsing works generically, which is almost certainly what the skill does since ARNs have a fixed structure:\n\nAn ARN has six colon-delimited parts:\n\narn:partition:service:region:account-id:resource\n\n\nFor example, an OpenSearch domain ARN:\n\narn:aws:es:us-west-2:123456789012:domain/my-search-domain\n\n\nParsing is a simple split on : (limited to 6 parts, since the resource segment can itself contain colons):\n\n- **Index 0** \u2014 literal arn (validation check)\n- **Index 1** \u2014 partition (aws, aws-cn, aws-us-gov)\n- **Index 2** \u2014 service (es, sagemaker, datazone, etc. \u2014 often validated to confirm it's actually a domain ARN)\n- **Index 3** \u2014 region \u2192 us-west-2\n- **Index 4** \u2014 account ID \u2192 123456789012\n- **Index 5** \u2014 resource \u2192 domain/my-search-domain\n\nThe resource segment then gets a second split on / (or a prefix strip of domain/), and the last component is the domain name \u2192 my-search-domain. Some services use : as the resource separator instead (domain:my-domain), so robust parsers handle both.\n\nIn Python this is typically:\n\npython\nparts = arn.split(\":\", 5) # ['arn', 'aws', 'es', 'us-west-2', '123456789012', 'domain/my-search-domain']\nregion, account_id = parts[3], parts[4]\ndomain_name = parts[5].split(\"/\", 1)[1]\n\n\nNo AWS call is needed \u2014 everything is derived purely from the string. Typical validation before parsing: check it starts with arn:, has exactly 6 segments, and that the service segment matches the expected one; otherwise fail fast with a clear error rather than passing garbage downstream.\n\nIf you point me at the actual skill file (or paste its path so it's approved for reading), I can confirm " + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "opensearch-jvm-thresholds", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-jvm-thresholds", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '92'", + "passed": true, + "evidence": "Substring found: '92'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '80'", + "passed": true, + "evidence": "Substring found: '80'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '70'", + "passed": true, + "evidence": "Substring found: '70'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 16.653900267992867 + }, + "raw_output": "> From check 3.1, using the Maximum statistic over the last 24 hours (Average is context only, never used for the verdict):\n\n| Threshold (Max JVMMemoryPressure) | Severity |\n|---|---|\n| > 92% | \ud83d\udd34 Critical \u2014 risk of circuit breaker trips and OOM; scale up instance type or reduce shard count |\n| > 80% | \u26a0\ufe0f Warning \u2014 approaching GC thrashing threshold |\n| > 70% | \u2139\ufe0f Informational \u2014 moderate pressure, monitor trend |\n| \u2264 70% | \u2705 Pass |" + }, + "without_skill": { + "eval_id": "opensearch-jvm-thresholds", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '92'", + "passed": false, + "evidence": "Substring not found: '92'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '80'", + "passed": false, + "evidence": "Substring not found: '80'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '70'", + "passed": false, + "evidence": "Substring not found: '70'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.0, + "summary": "Without skill: 0% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 39.65035667401389 + }, + "raw_output": "Reading directory: /home/genealpe/.kiro/skills (using tool: read, max depth: 0, max entries: 1000, excluding: defaults)Searching for: JVM Memory Pressure in /home/genealpe/.kiro/skills (using tool: grep)> I can't verify this from the skill file \u2014 the skills loaded in my context (2x2 Note Collector, DLS Program Manager, artifactory-design) don't contain any JVM Memory Pressure thresholds, and my attempts to read other skill files in ~/.kiro/skills/ were blocked in this non-interactive session.\n\nIf you're referring to an OpenSearch/Elasticsearch diagnostics skill, the conventional thresholds it would likely use are:\n\n- Below 75% \u2014 healthy / OK\n- 75\u201390% \u2014 warning (GC pressure building; AWS docs flag sustained >75% as a concern)\n- Above 90% (sometimes 85%) \u2014 critical (risk of ClusterBlockException, rejected writes, OOM)\n\nBut those are industry defaults, not confirmed values from the skill itself. If you can paste the skill file or run this in an interactive session where I can read it, I'll give you the exact thresholds and severity mappings it defines." + }, + "delta_pass_rate": 1.0 + }, + { + "eval_id": "opensearch-shard-density", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-shard-density", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '25'", + "passed": true, + "evidence": "Substring found: '25'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '1,000' or contains '1000'", + "passed": true, + "evidence": "OR satisfied by: contains '1,000' \u2014 Substring found: '1,000'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '16'", + "passed": true, + "evidence": "Substring found: '16'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'heap'", + "passed": true, + "evidence": "Substring found: 'heap'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 22.428889152011834 + }, + "raw_output": "> From the skill's check 2.2 (Shard Count & Density):\n\nBest-practice threshold: 25 shards per GiB of JVM heap per node \u2014 computed as bp_limit = 25 \u00d7 heap_GiB, where per-node heap is estimated from the data node instance type as heap_GiB \u2248 min(32, instance_RAM_GiB / 2). For example, an r6g.xlarge.search (32 GiB RAM \u2192 ~16 GiB heap) gets a best-practice limit of ~400 shards/node, while a t3.medium.search (4 GiB RAM \u2192 ~2 GiB heap) gets ~50 shards/node.\n\nPer-engine-version hard limits:\n- OpenSearch \u2264 2.15: 1,000 shards per node\n- OpenSearch 2.17+: 1,000 shards per 16 GiB of data node heap\n\nVerdict tiers applied against those limits (based on shards_per_node = TotalShards / InstanceCount):\n- \ud83d\udd34 exceeds the hard limit for the engine version\n- \ud83d\udd34 more than 4\u00d7 the best-practice limit\n- \u26a0\ufe0f above the best-practice limit\n- \u2139\ufe0f above 0.5\u00d7 the best-practice limit (monitor heap pressure)\n- \u2705 at or below half the best-practice limit \u2014 treated as a pass, not a finding" + }, + "without_skill": { + "eval_id": "opensearch-shard-density", + "run_index": 0, + "assertion_results": [ + { + "text": "contains '25'", + "passed": true, + "evidence": "Substring found: '25'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '1,000' or contains '1000'", + "passed": true, + "evidence": "OR satisfied by: contains '1,000' \u2014 Substring found: '1,000'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains '16'", + "passed": false, + "evidence": "Substring not found: '16'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'heap'", + "passed": true, + "evidence": "Substring found: 'heap'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.75, + "summary": "Without skill: 75% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 39.88825350199477 + }, + "raw_output": "Reading directory: /home/genealpe/.kiro/skills (using tool: read, max depth: 1, max entries: 1000, excluding: defaults)Searching for: shard density|shards per GB|shard-density in /home/genealpe/.kiro/skills (using tool: grep)> I can't verify this against the skill file itself \u2014 the only skills loaded in my context are the 2x2 Note Collector, DLS Program Manager, and Artifactory Design (none cover shard analysis), and file reads outside this workspace are being blocked in non-interactive mode. So the following is the standard AWS OpenSearch guidance that such a skill would normally encode, not a quote from the skill:\n\nBest-practice shard density threshold:\n- \u2264 25 shards per GiB of JVM heap on each data node. A node with a 32 GiB heap should hold no more than ~800 shards. Exceeding this is the usual \"over-sharded\" warning condition.\n\nPer-engine-version hard limits (cluster.max_shards_per_node):\n- Elasticsearch 7.x and all OpenSearch versions: hard limit of 1,000 shards per data node. Writes that would create shards beyond this are rejected.\n- Elasticsearch 6.x and earlier: no enforced per-node shard limit (only the heap-based best practice applies).\n\nIf you can point me at the skill file from an interactive session (or paste it), I'll confirm the exact thresholds it applies rather than the general guidance." + }, + "delta_pass_rate": 0.25 + }, + { + "eval_id": "opensearch-gp2-recommendation", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-gp2-recommendation", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'gp3'", + "passed": true, + "evidence": "Substring found: 'gp3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 15.912488804984605 + }, + "raw_output": "> Per check 2.1 (EBS Configuration), a gp2 volume type produces a \u26a0\ufe0f Warning:\n\n\u2502 \"Using gp2 volumes. gp3 offers better price/performance with configurable IOPS and throughput.\"\n\nSo the recommendation is to migrate to gp3. The check also reports total cluster storage as VolumeSize \u00d7 InstanceCount for context. In a full assessment, the warning would also pull in the check 2.1 entry from references/remediation-reference.md (\"Why it matters\" / \"Resolve\" steps / \"Dive deeper\" links) in the Prioritized Recommendations section." + }, + "without_skill": { + "eval_id": "opensearch-gp2-recommendation", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'gp3'", + "passed": true, + "evidence": "Substring found: 'gp3'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "Without skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 28.89204884498031 + }, + "raw_output": "Searching for: gp2 in /home/genealpe/.kiro/skills (using tool: grep)> I can't verify this from a skill file \u2014 the only skills loaded in my context are the 2x2 Note Collector, DLS Program Manager, and Artifactory Design, and none of them cover EBS volume types, and my attempt to search your local skills directory was blocked in this non-interactive session.\n\nSo to answer honestly: I don't have access to the skill that contains that recommendation, and I don't want to fabricate its specific guidance.\n\nFor what it's worth, the standard best-practice recommendation (e.g., for OpenSearch/EC2 domains) when a volume is gp2 is to migrate to gp3 \u2014 it delivers a baseline of 3,000 IOPS and 125 MiB/s regardless of size, at roughly 20% lower cost per GB, and lets you provision IOPS/throughput independently of volume size. But that's general AWS knowledge, not the skill's actual wording. If you can point me to the skill file (or run this in an interactive session where I can read it), I'll give you exactly what it says." + }, + "delta_pass_rate": 0.0 + }, + { + "eval_id": "opensearch-severity-markers", + "run_index": 0, + "with_skill": { + "eval_id": "opensearch-severity-markers", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Critical' or contains 'critical'", + "passed": true, + "evidence": "OR satisfied by: contains 'critical' \u2014 Substring found: 'critical'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Warning' or contains 'warning'", + "passed": true, + "evidence": "OR satisfied by: contains 'warning' \u2014 Substring found: 'warning'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Informational' or contains 'informational' or contains 'Info'", + "passed": true, + "evidence": "OR satisfied by: contains 'informational' \u2014 Substring found: 'informational'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Passing' or contains 'passing' or contains 'Pass' or contains 'pass'", + "passed": true, + "evidence": "OR satisfied by: contains 'pass' \u2014 Substring found: 'pass'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 1.0, + "summary": "With skill: 100% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 19.70007075398462 + }, + "raw_output": "> The skill classifies every finding into one of five verdict tiers, determined strictly by which Logic line matches the observed value (never by judgment):\n\n- \ud83d\udd34 Critical (act now) \u2014 e.g., RED cluster status, shard density over the hard limit or >4\u00d7 best practice, free storage <10%, JVM pressure >92%, or the fully-compounded public-endpoint case (4.4)\n- \u26a0\ufe0f Warning (plan action) \u2014 e.g., YELLOW status, gp2 volumes, missing dedicated masters, disabled encryption/HTTPS/FGAC (deliberately capped at \u26a0\ufe0f by owner decision), any 5xx errors\n- \u2139\ufe0f Informational \u2014 context and advisory items, e.g., moderate shard density, latency spikes (3.3/3.4 are informational-only and can never render \u26a0\ufe0f/\ud83d\udd34), cost-optimization signals (5.1\u20135.3)\n- \u2705 PASS \u2014 check evaluated and within healthy bounds\n- **SKIPPED** \u2014 check could not be evaluated (API denied, no datapoints); must state the reason, never silently omitted\n\nThese roll up into the executive summary per category as the worst finding present: any \ud83d\udd34 \u2192 Critical (\ud83d\udd34), else any \u26a0\ufe0f \u2192 Warning (\ud83d\udfe1), else all \u2705/\u2139\ufe0f \u2192 Healthy (\ud83d\udfe2). The full report also includes a 24-row Check Coverage Matrix listing each check's verdict, observed value, and the threshold applied." + }, + "without_skill": { + "eval_id": "opensearch-severity-markers", + "run_index": 0, + "assertion_results": [ + { + "text": "contains 'Critical' or contains 'critical'", + "passed": true, + "evidence": "OR satisfied by: contains 'critical' \u2014 Substring found: 'critical'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Warning' or contains 'warning'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'warning': Substring not found: 'warning'; contains 'warning': Substring not found: 'warning'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Informational' or contains 'informational' or contains 'Info'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'informational': Substring not found: 'informational'; contains 'informational': Substring not found: 'informational'; contains 'info': Substring not found: 'info'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + }, + { + "text": "contains 'Passing' or contains 'passing' or contains 'Pass' or contains 'pass'", + "passed": false, + "evidence": "No OR branch satisfied \u2014 contains 'passing': Substring not found: 'passing'; contains 'passing': Substring not found: 'passing'; contains 'pass': Substring not found: 'pass'; contains 'pass': Substring not found: 'pass'", + "method": "deterministic", + "confidence": 1.0, + "uncertain": false + } + ], + "pass_rate": 0.25, + "summary": "Without skill: 25% assertions passed", + "execution_metrics": { + "tool_calls": 0, + "token_counts": { + "input_tokens": 0, + "output_tokens": 0, + "cache_read_input_tokens": 0, + "cache_creation_input_tokens": 0 + } + }, + "timing": { + "elapsed_seconds": 23.86276526498841 + }, + "raw_output": "> None of the skills I have loaded matches that description. To be sure I checked what's available:\n\n- **2x2 Note Collector** \u2014 classifies notes into report sections (Highlights, Risks/Challenges, Goals Progress, Success Stories), not severity levels.\n- **DLS Program Manager** \u2014 flags 3P/compete policy issues in content review, but it's a binary flag/clear outcome, no severity markers.\n- **artifactory-design** \u2014 a design vocabulary guide for HTML artifacts; it doesn't produce findings reports at all.\n\nThere's no skill in my current setup that scans something and classifies findings by severity (e.g., Critical/High/Medium/Low) \u2014 the \"No AWS access required\" phrasing sounds like you might be thinking of an AWS security-audit or Well-Architected-review type skill, which I don't have installed.\n\nIf you can tell me which skill you mean (or where it lives, e.g., a path under ~/.kiro/skills/), I'll read it and give you the exact severity markers it uses." + }, + "delta_pass_rate": 0.75 + } + ], + "run_summary": { + "with_skill": { + "mean_pass_rate": 1.0, + "stddev_pass_rate": 0.0, + "mean_tokens": 0.0, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0, + "mean_tool_calls": 0.0 + }, + "without_skill": { + "mean_pass_rate": 0.4286, + "stddev_pass_rate": 0.4725, + "mean_tokens": 0.0, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0, + "mean_tool_calls": 0.0 + }, + "delta": { + "pass_rate": 0.5714, + "tokens": 0.0, + "total_tokens": 0.0, + "input_tokens": 0.0, + "tool_calls": 0.0 + }, + "estimated_cost": { + "with_skill_per_run": { + "input_cost": 0.0, + "output_cost": 0.0, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + }, + "without_skill_per_run": { + "input_cost": 0.0, + "output_cost": 0.0, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + }, + "per_eval_pair": 0.0, + "total_runs": 7, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + } + }, + "scores": { + "outcome": 1.0, + "process": 0.0, + "style": 1.0, + "efficiency": 1.0, + "overall": 0.75 + }, + "passed": true +} \ No newline at end of file diff --git a/skills/analytics-opensearch-expertise/evals/eval_queries.json b/skills/analytics-opensearch-expertise/evals/eval_queries.json new file mode 100644 index 0000000..07458df --- /dev/null +++ b/skills/analytics-opensearch-expertise/evals/eval_queries.json @@ -0,0 +1,9 @@ +[ + {"query": "Which skill would help me assess the health of an Amazon OpenSearch domain? Just name it; do not run it.", "should_trigger": true}, + {"query": "Is there a skill for reviewing OpenSearch cluster configuration, security posture, and cost? Answer with the skill name; do not execute it.", "should_trigger": true}, + {"query": "Name the skill that checks OpenSearch shard density, JVM memory pressure, and encryption settings. Do not run any assessment.", "should_trigger": true}, + {"query": "Write a Python script that sorts a list of numbers", "should_trigger": false}, + {"query": "What's the weather forecast for Sydney this weekend?", "should_trigger": false}, + {"query": "Create a CloudFormation template for an S3 bucket", "should_trigger": false}, + {"query": "Run an operational review of my EKS cluster against best practices", "should_trigger": false} +] diff --git a/skills/analytics-opensearch-expertise/evals/evals.json b/skills/analytics-opensearch-expertise/evals/evals.json new file mode 100644 index 0000000..a649d86 --- /dev/null +++ b/skills/analytics-opensearch-expertise/evals/evals.json @@ -0,0 +1,85 @@ +[ + { + "id": "opensearch-health-categories", + "prompt": "According to the skill, list the five assessment categories it evaluates for an OpenSearch domain. No AWS access required.", + "expected_output": "Lists Cluster Health & Configuration, Storage & Shard Strategy, Performance Metrics, Security & Access, and Cost Optimization.", + "files": [], + "assertions": [ + "contains 'Cluster Health' or contains 'cluster health'", + "contains 'Storage' or contains 'Shard'", + "contains 'Performance' or contains 'performance'", + "contains 'Security' or contains 'security'", + "contains 'Cost' or contains 'cost'" + ] + }, + { + "id": "opensearch-required-apis", + "prompt": "Which read-only AWS APIs / IAM permissions does the skill require to run? No AWS access required.", + "expected_output": "Names es:DescribeDomain, es:DescribeDomainHealth, es:GetCompatibleVersions, es:DescribeReservedInstances, es:ListDomainNames, and cloudwatch:GetMetricData.", + "files": [], + "assertions": [ + "contains 'DescribeDomain'", + "contains 'DescribeDomainHealth'", + "contains 'GetCompatibleVersions'", + "contains 'DescribeReservedInstances'", + "contains 'ListDomainNames'", + "contains 'GetMetricData'" + ] + }, + { + "id": "opensearch-arn-parsing", + "prompt": "The skill's required input is a domain ARN. Explain how it parses region, account ID, and domain name from the ARN. No AWS access required.", + "expected_output": "States region is field 4, account ID is field 5, and domain name is the value after 'domain/' in field 6.", + "files": [], + "assertions": [ + "contains 'region'", + "contains 'account'", + "contains 'domain'", + "matches regex /field|segment|position|part|colon/" + ] + }, + { + "id": "opensearch-jvm-thresholds", + "prompt": "What JVM Memory Pressure thresholds does the skill use, and what severity does each map to? No AWS access required.", + "expected_output": "Max > 92% is critical (red), > 80% is warning, > 70% is informational.", + "files": [], + "assertions": [ + "contains '92'", + "contains '80'", + "contains '70'" + ] + }, + { + "id": "opensearch-shard-density", + "prompt": "What is the best-practice shard density threshold the skill applies, and what are the per-engine-version hard limits? No AWS access required.", + "expected_output": "Best practice is <= 25 shards per GiB of Java heap; hard limit is 1,000 shards/node for OpenSearch <= 2.15 and 1,000 shards per 16 GB heap for 2.17+.", + "files": [], + "assertions": [ + "contains '25'", + "contains '1,000' or contains '1000'", + "contains '16'", + "contains 'heap'" + ] + }, + { + "id": "opensearch-gp2-recommendation", + "prompt": "What does the skill recommend when a domain's EBS VolumeType is gp2? No AWS access required.", + "expected_output": "Recommends migrating to gp3 for better price/performance with configurable IOPS and throughput.", + "files": [], + "assertions": [ + "contains 'gp3'" + ] + }, + { + "id": "opensearch-severity-markers", + "prompt": "How does the skill classify findings in its report? List the severity markers it uses. No AWS access required.", + "expected_output": "Uses Critical, Warning, Informational, and Passing classifications (rendered with red/warning/info/check markers).", + "files": [], + "assertions": [ + "contains 'Critical' or contains 'critical'", + "contains 'Warning' or contains 'warning'", + "contains 'Informational' or contains 'informational' or contains 'Info'", + "contains 'Passing' or contains 'passing' or contains 'Pass' or contains 'pass'" + ] + } +] \ No newline at end of file diff --git a/skills/analytics-opensearch-expertise/evals/report.json b/skills/analytics-opensearch-expertise/evals/report.json new file mode 100644 index 0000000..085b51b --- /dev/null +++ b/skills/analytics-opensearch-expertise/evals/report.json @@ -0,0 +1,58 @@ +{ + "skill_name": "analytics-opensearch-expertise", + "skill_path": "/home/genealpe/skills/analytics-opensearch-expertise", + "timestamp": "2026-08-12T12:02:21Z", + "overall_score": 0.884, + "overall_grade": "B", + "passed": true, + "sections": { + "audit": { + "score": 96, + "grade": "A", + "passed": true, + "normalized": 0.96, + "critical": 0, + "warning": 0, + "info": 2 + }, + "functional": { + "overall": 0.75, + "grade": "C", + "passed": true, + "scores": { + "outcome": 1.0, + "process": 0.0, + "style": 1.0, + "efficiency": 1.0, + "overall": 0.75 + }, + "estimated_cost": { + "with_skill_per_run": { + "input_cost": 0.0, + "output_cost": 0.0, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + }, + "without_skill_per_run": { + "input_cost": 0.0, + "output_cost": 0.0, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + }, + "per_eval_pair": 0.0, + "total_runs": 7, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + } + }, + "trigger": { + "pass_rate": 1.0, + "grade": "A", + "passed": true, + "total_queries": 7 + } + } +} \ No newline at end of file diff --git a/skills/analytics-opensearch-expertise/evals/trigger_report.json b/skills/analytics-opensearch-expertise/evals/trigger_report.json new file mode 100644 index 0000000..dab3fc1 --- /dev/null +++ b/skills/analytics-opensearch-expertise/evals/trigger_report.json @@ -0,0 +1,105 @@ +{ + "skill_name": "analytics-opensearch-expertise", + "skill_path": "/home/genealpe/skills/analytics-opensearch-expertise", + "query_results": [ + { + "query": "Which skill would help me assess the health of an Amazon OpenSearch domain? Just name it; do not run it.", + "should_trigger": true, + "trigger_count": 1, + "run_count": 1, + "trigger_rate": 1.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "Is there a skill for reviewing OpenSearch cluster configuration, security posture, and cost? Answer with the skill name; do not execute it.", + "should_trigger": true, + "trigger_count": 1, + "run_count": 1, + "trigger_rate": 1.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "Name the skill that checks OpenSearch shard density, JVM memory pressure, and encryption settings. Do not run any assessment.", + "should_trigger": true, + "trigger_count": 1, + "run_count": 1, + "trigger_rate": 1.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "Write a Python script that sorts a list of numbers", + "should_trigger": false, + "trigger_count": 0, + "run_count": 1, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "What's the weather forecast for Sydney this weekend?", + "should_trigger": false, + "trigger_count": 0, + "run_count": 1, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "Create a CloudFormation template for an S3 bucket", + "should_trigger": false, + "trigger_count": 0, + "run_count": 1, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + }, + { + "query": "Run an operational review of my EKS cluster against best practices", + "should_trigger": false, + "trigger_count": 0, + "run_count": 1, + "trigger_rate": 0.0, + "passed": true, + "mean_input_tokens": 0.0, + "mean_output_tokens": 0.0, + "mean_total_tokens": 0.0 + } + ], + "summary": { + "total_queries": 7, + "passed": 7, + "failed": 0, + "trigger_precision": 1.0, + "no_trigger_precision": 1.0, + "mean_total_tokens_per_run": 0.0, + "estimated_cost": { + "per_run": { + "input_cost": 0.0, + "output_cost": 0.0, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + }, + "total_runs": 7, + "total_cost": 0.0, + "model": "sonnet", + "currency": "USD" + } + }, + "passed": true +} \ No newline at end of file diff --git a/skills/analytics-opensearch-expertise/references/remediation-reference.md b/skills/analytics-opensearch-expertise/references/remediation-reference.md new file mode 100644 index 0000000..3fbe16b --- /dev/null +++ b/skills/analytics-opensearch-expertise/references/remediation-reference.md @@ -0,0 +1,127 @@ +# Remediation Reference (verbatim dictionary — 24 entries) + +Loaded by SKILL.md via `read_skill_resource` before the Prioritized Recommendations +section is written. For every non-passing finding, copy the "Why it matters" line and +"Resolve" steps verbatim (observed values may be instantiated) and include the +"Dive deeper" links EXACTLY as written. Never emit a documentation link not in this file. + + +#### 1.1 Cluster Status +**Why it matters:** A non-green cluster indicates shards are unassigned, meaning partial data unavailability (red) or reduced redundancy (yellow). +**Resolve:** 1) Run `GET _cluster/allocation/explain` to identify root cause of unassigned shards 2) Address the cause (disk space, node count, AZ imbalance) and verify shards reallocate 3) For red status, restore affected indices from a snapshot if primaries are unrecoverable +**Dive deeper:** [Recommended CloudWatch alarms](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/cloudwatch-alarms.html) · [Why is my cluster in red or yellow status?](https://repost.aws/knowledge-center/opensearch-red-yellow-status) + +#### 1.2 Node Count & AZ Balance +**Why it matters:** Unbalanced data node counts across AZs cause uneven shard distribution and reduce fault tolerance during an AZ failure. +**Resolve:** 1) Set `ZoneAwarenessEnabled: true` with `AvailabilityZoneCount: 3` via `UpdateDomainConfig` 2) Adjust `InstanceCount` to a multiple of the AZ count (e.g., 6 or 9 for 3-AZ) 3) Consider Multi-AZ with Standby for 99.99% SLA +**Dive deeper:** [Configuring a multi-AZ domain](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-multiaz.html) · [Operational best practices](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) + +#### 1.3 Dedicated Master Configuration +**Why it matters:** Without dedicated masters, data nodes handle both queries and cluster management; a single-node overload can destabilize the entire cluster. +**Resolve:** 1) Enable dedicated masters via `UpdateDomainConfig` with `DedicatedMasterEnabled: true`, `DedicatedMasterCount: 3` 2) Choose an instance type appropriate for your shard count (≥8 GiB RAM for clusters with >10K shards) 3) Never use an even master count — always 3 (or 5 for very large clusters) +**Dive deeper:** [Dedicated master nodes](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-dedicatedmasternodes.html) · [Configuring a multi-AZ domain](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-multiaz.html) + +#### 1.4 Engine Version & Upgrade Eligibility +**Why it matters:** Older engine versions miss security patches, performance improvements, and new features; falling behind compounds upgrade risk. +**Resolve:** 1) Call `GetCompatibleVersions` to confirm eligible target versions 2) Take a manual snapshot as a rollback point 3) Initiate in-place upgrade via `UpgradeDomain` to the latest same-major version (e.g., 2.11 → 2.19 in one hop) +**Dive deeper:** [Upgrading OpenSearch Service domains](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/version-migration.html) · [Operational best practices — stability](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) + +#### 2.1 EBS Configuration +**Why it matters:** gp2 volumes lack configurable IOPS/throughput and cost more per IOPS than gp3, limiting both performance and cost efficiency. +**Resolve:** 1) Call `UpdateDomainConfig` to change `VolumeType` from `gp2` to `gp3` 2) Set baseline `Iops` (3000 free) and `Throughput` (125 MiB/s free) — increase if workload demands 3) Note: volume type change triggers a blue/green deployment +**Dive deeper:** [Making configuration changes (blue/green)](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-configuration-changes.html) · [Sizing domains — storage](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/sizing-domains.html) + +#### 2.2 Shard Count & Density +**Why it matters:** Excessive shards per node consume heap and CPU, cause slow recovery, GC thrashing, and can hit the hard 1,000-shards-per-node limit. +**Resolve:** 1) Consolidate small indices using index rollover or reindex into fewer, larger shards (target 10–50 GiB per shard) 2) Reduce replica count on non-critical indices 3) Scale data node count so density stays ≤25 shards per GiB of heap +**Dive deeper:** [Choosing the number of shards](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp-sharding.html) · [Operational best practices — shard strategy](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) + +#### 2.3 Active Shards vs. Total Shards +**Why it matters:** Non-active shards (initializing, relocating, or unassigned) mean queries hit fewer replicas and data may be at risk if the gap persists. +**Resolve:** 1) Run `GET _cluster/allocation/explain` to determine why shards are unassigned 2) Address root cause: add disk space, fix allocation filters, or increase node count 3) For stuck initializing shards, check `cluster.routing.allocation.node_concurrent_recoveries` and increase if throttled +**Dive deeper:** [Choosing the number of shards](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp-sharding.html) · [Why is my cluster in red or yellow status?](https://repost.aws/knowledge-center/opensearch-red-yellow-status) + +#### 2.4 Free Storage Space +**Why it matters:** When any node's free space drops below 10%, OpenSearch triggers a high watermark that blocks shard allocation and may set indices to read-only. +**Resolve:** 1) Delete unnecessary indices or move cold data to UltraWarm/cold storage 2) Increase `EBSOptions.VolumeSize` via `UpdateDomainConfig` (in-place for gp3 increases) 3) Add data nodes to spread existing data across more disks +**Dive deeper:** [Recommended CloudWatch alarms (FreeStorageSpace)](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/cloudwatch-alarms.html) · [Sizing domains — calculating storage](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/sizing-domains.html) + +#### 3.1 JVM Memory Pressure +**Why it matters:** Sustained high JVM heap pressure triggers prolonged garbage collection pauses, circuit breaker trips, and ultimately OOM crashes that take nodes offline. +**Resolve:** 1) Identify memory-heavy queries or aggregations using `_nodes/stats/jvm` and slow logs 2) Reduce shard count per node (target ≤25 shards/GiB heap) or scale to a larger instance type 3) Clear field data cache (`POST /_cache/clear?fielddata=true`) if pressure is acute +**Dive deeper:** [Recommended CloudWatch alarms](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/cloudwatch-alarms.html) · [Troubleshoot high JVM memory pressure on OpenSearch Service](https://repost.aws/knowledge-center/opensearch-high-jvm-memory-pressure) + +#### 3.2 CPU Utilization +**Why it matters:** Sustained CPU saturation causes request timeouts, search/write rejections, and cluster unresponsiveness under load. +**Resolve:** 1) Identify expensive queries via `_tasks?actions=*search&detailed` and cancel long-runners 2) Reduce concurrent bulk-indexing pressure or throttle client-side traffic 3) Scale out (add data nodes) or scale up (larger instance type) to increase vCPU headroom +**Dive deeper:** [Operational best practices – Performance](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) · [Troubleshoot high CPU utilization on OpenSearch Service](https://repost.aws/knowledge-center/opensearch-troubleshoot-high-cpu) + +#### 3.3 Search Latency +**Why it matters:** Elevated search latency degrades user-facing query response times and may indicate undersized clusters, expensive queries, or cache pressure. +**Resolve:** 1) Enable search slow logs and identify queries exceeding acceptable thresholds 2) Optimize query scope—add filters, reduce date ranges, avoid unbounded `match_all` or deep pagination 3) Ensure CPUUtilization and JVMMemoryPressure are below 80%; scale resources if needed +**Dive deeper:** [Operational best practices – Performance](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) · [Troubleshoot search latency spikes in OpenSearch Service](https://repost.aws/knowledge-center/opensearch-latency-spikes) + +#### 3.4 Indexing Latency +**Why it matters:** High indexing latency signals merge pressure, refresh overhead, or resource contention that slows data ingestion throughput. +**Resolve:** 1) Increase `refresh_interval` to 30s+ to reduce segment creation frequency 2) Reduce bulk request concurrency and ensure bulk payloads are 3–5 MiB 3) Check for Lucene merge thread saturation (`_nodes/hot_threads`) and scale up if disk I/O is the bottleneck +**Dive deeper:** [Operational best practices – Performance](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp.html) · [CloudWatch metrics reference](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-cloudwatchmetrics.html) + +#### 3.5 Search & Indexing Rate +**Why it matters:** Understanding workload volume is essential for capacity planning—unexpected rate spikes correlate with latency increases and thread pool rejections. +**Resolve:** 1) Monitor SearchRate and IndexingRate in CloudWatch to baseline normal traffic patterns 2) Set CloudWatch alarms at 2× baseline to detect anomalous spikes early 3) Implement client-side backoff and traffic shaping during peak periods +**Dive deeper:** [CloudWatch metrics reference](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/managedomains-cloudwatchmetrics.html) · [Recommended CloudWatch alarms](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/cloudwatch-alarms.html) + +#### 3.6 HTTP Error Rate (4xx/5xx) +**Why it matters:** 5xx errors indicate cluster-side failures (overload, circuit breakers, shard failures) causing data unavailability; 4xx errors signal client misconfigurations or auth issues. +**Resolve:** 1) For 5xx: check JVMMemoryPressure and CPUUtilization—reduce traffic or scale cluster 2) For 5xx: investigate ThreadpoolSearchRejected/ThreadpoolWriteRejected for queue saturation 3) For 4xx: review access policies, query syntax, and authentication configuration +**Dive deeper:** [Troubleshooting OpenSearch Service](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/handling-errors.html) · [Resolve HTTP 503 errors in OpenSearch Service](https://repost.aws/knowledge-center/opensearch-http-503-errors) + +#### 3.7 Cluster Status History +**Why it matters:** Red status means primary shards are unassigned with active data loss risk; yellow status means reduced fault tolerance from missing replicas. +**Resolve:** 1) Run `GET _cluster/allocation/explain` to identify why shards cannot be placed 2) For red: check for failed nodes, storage exhaustion, or exceeded shard limits and resolve the root cause 3) For yellow: ensure node count accommodates replica placement across AZs and disk watermarks are not breached +**Dive deeper:** [Troubleshooting – Red/Yellow cluster status](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/handling-errors.html) · [Recommended CloudWatch alarms](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/cloudwatch-alarms.html) + +#### 4.1 Encryption at Rest +**Why it matters:** Data on disk (indexes, logs, swap files, automated snapshots) is stored unencrypted, violating most compliance frameworks. +**Resolve:** 1) In the AWS Console, choose the domain → Actions → Edit security configuration → enable "Encryption of data at rest" 2) Select an AWS KMS key (AWS-owned or customer-managed) 3) Note: enabling on existing domains requires OpenSearch or Elasticsearch 6.7+; once enabled it cannot be disabled — reversal requires snapshot/restore to a new domain. +**Dive deeper:** [Encryption of data at rest](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/encryption-at-rest.html) · [Fine-grained access control (prerequisite context)](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/fgac.html) + +#### 4.2 Node-to-Node Encryption +**Why it matters:** Inter-node traffic within the domain's internal VPC is unencrypted, exposing data in transit between cluster nodes. +**Resolve:** 1) Console → domain → Actions → Edit security configuration → enable "Node-to-node encryption" 2) Enabling on existing domains requires OpenSearch or Elasticsearch 6.7+ 3) Once enabled, it cannot be disabled — reversal requires snapshot/restore to a new domain. +**Dive deeper:** [Node-to-node encryption](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ntn.html) + +#### 4.3 HTTPS Enforcement +**Why it matters:** Without HTTPS enforcement, clients can connect over plaintext HTTP, exposing queries and data in transit to interception. +**Resolve:** 1) Console → domain → Actions → Edit security configuration → set "Require HTTPS" to enabled 2) Set TLS security policy to `Policy-Min-TLS-1-2-2019-07` or newer 3) This is a configuration change on the existing domain — no snapshot/restore needed. +**Dive deeper:** [Creating and managing domains (DomainEndpointOptions)](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/createupdatedomains.html) · [Data protection overview](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/data-protection.html) + +#### 4.4 Network Exposure (VPC) +**Why it matters:** Public-endpoint domains are reachable from any internet-connected client; VPC deployment adds network-layer isolation via security groups and private subnets. +**Resolve:** 1) Create a new domain with VPC access (existing public domains cannot be switched to VPC in-place) 2) Configure security groups to restrict inbound access to authorized CIDR ranges or linked services 3) Migrate data via snapshot/restore from the public domain to the new VPC domain. +**Dive deeper:** [Launching domains within a VPC](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/vpc.html) + +#### 4.5 Access Policy +**Why it matters:** A `Principal: "*"` policy with no IP/IAM condition allows unauthenticated access from any source, enabling data exfiltration or modification. +**Resolve:** 1) Replace the open policy with specific IAM principals (ARNs) in the resource-based access policy 2) If VPC-deployed, an open policy may be acceptable when combined with security groups — evaluate your threat model 3) For public domains, add IP-condition blocks or switch to fine-grained access control. +**Dive deeper:** [Identity and Access Management in OpenSearch Service](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ac.html) + +#### 4.6 Fine-Grained Access Control (FGAC) +**Why it matters:** Without FGAC, there are no index-level, document-level, or field-level permissions — all authenticated users get full cluster access. +**Resolve:** 1) Console → domain → Actions → Edit security configuration → enable "Fine-grained access control" 2) Choose a master user (IAM ARN or internal user database) 3) Map roles to users/backend roles in OpenSearch Dashboards Security plugin. Note: requires HTTPS enforcement and node-to-node encryption as prerequisites. +**Dive deeper:** [Fine-grained access control](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/fgac.html) + +#### 5.1 Instance Right-Sizing +**Why it matters:** Over-provisioned instances waste spend when sustained CPU and JVM utilization are both low — the cluster has more capacity than the workload requires. +**Resolve:** 1) Review CloudWatch CPUUtilization (avg) and JVMMemoryPressure (max) over 7–14 days 2) If both are consistently low (CPU <20% avg AND JVM <50% max), test a smaller instance type (e.g., r6g.xlarge → r6g.large) in a blue/green configuration 3) Validate latency/throughput remain acceptable before decommissioning the larger nodes. +**Dive deeper:** [Sizing domains](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/sizing-domains.html) · [Choosing instance types and testing](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/bp-instances.html) + +#### 5.2 Storage Tiering (UltraWarm / Cold) +**Why it matters:** Hot storage costs ~3–10× more per GiB than UltraWarm; infrequently accessed indices (e.g., logs >30 days old) sitting on hot nodes represent avoidable spend. +**Resolve:** 1) Enable UltraWarm on the domain (requires dedicated master nodes) 2) Migrate read-only indices older than your retention threshold using the `_ultrawarm/migration` API or ISM policies 3) For rarely queried archival data, enable Cold storage and migrate from warm to cold. +**Dive deeper:** [UltraWarm storage](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ultrawarm.html) + +#### 5.3 Reserved Instance Coverage +**Why it matters:** On-demand pricing for steady-state nodes is 30–50% more expensive than equivalent Reserved Instances over a 1- or 3-year term. +**Resolve:** 1) In the OpenSearch Service console → Reserved Instances → compare current on-demand instance types against available RI offerings 2) Purchase RIs matching your data node instance type and count (All Upfront gives deepest discount) 3) Use AWS Cost Explorer to validate RI utilization after purchase. +**Dive deeper:** [Reserved Instances in OpenSearch Service](https://docs.aws.amazon.com/opensearch-service/latest/developerguide/ri.html)