概要 #
CPU使用率の異常にコンソールを開いて毎回気づくのは現実的ではありません。しきい値を決めて自動で通知させたいところです。
Terraformでは、Alert PolicyとNotification Channelをコードとして管理できます。ここではGCE VMのCPU使用率が一定時間Thresholdを超えた場合を条件にします。任意でEmail Notification Channelも接続します。
Terraformの基本操作は前提としています。Cloud Monitoringを初めて扱う方向けです。
このサンプル自体はVMを作成しないため、対象VMがなければAlertは通常発火しません。ダッシュボードの作成やLog-based Metricsは扱いません。
検証すること #
- GCE CPU使用率を監視するAlert Policyを作成できること
- Email Notification Channelを任意で追加できること
- Threshold、集計期間、継続時間を変数で調整できること
前提環境 #
- Google Cloud CLIとApplication Default Credentials(ADC)
- Terraform 1.5以降
- Cloud Monitoringを操作できる検証用Google Cloud Project
- Email通知を試す場合は受信可能なEmail Address
- 検証時のバージョン:Terraform 1.14.3、google 7.43.0
今回の構成 #
Alert PolicyとNotification ChannelはProject単位のResourceです。監視対象のVMとは別の場所にあります。
flowchart TB
Mail["メール受信者"]
subgraph GCP["Google Cloud"]
subgraph Project["Project"]
subgraph Mon["Cloud Monitoring"]
Policy["Alert Policy
tf-example-gce-cpu-high"]
Channel["Notification Channel
type: email"]
Policy -->|"通知先"| Channel
end
subgraph Region["asia-northeast1"]
VM["GCE Instance
(監視対象)"]
end
end
end
VM -.->|"メトリクスを送信"| Policy
Channel -->|"アラートメール"| Mail
点線がメトリクスの収集、実線が通知の経路です。監視対象のVMがなくてもPolicy自体は作成できます。
しきい値を超えたときの流れは次のとおりです。
sequenceDiagram
participant VM as GCE Instance
participant Mon as Cloud Monitoring
participant Policy as Alert Policy
participant Ch as Notification Channel
actor User as メール受信者
VM->>Mon: CPU使用率を継続的に送信
Mon->>Policy: 条件を評価
Note over Policy: CPU > 0.8 が 300秒継続
Policy->>Ch: インシデントを通知
Ch->>User: アラートメールを送信
使用するTerraformコード #
任意のEmail Channel #
resource "google_monitoring_notification_channel" "email" {
count = var.notification_email == null ? 0 : 1
display_name = var.notification_channel_name
type = "email"
enabled = true
labels = {
email_address = coalesce(var.notification_email, "noreply@example.com")
}
}
notification_emailのDefaultはnullです。未指定ならcount = 0となり、Alert Policyだけを作成します。指定時だけEmail Channelを作れるため、公開サンプルへ実Addressを埋め込む必要がありません。
CPU使用率のAlert Policy #
resource "google_monitoring_alert_policy" "cpu_usage" {
display_name = var.alert_policy_name
combiner = "OR"
enabled = true
conditions {
display_name = "GCE CPU utilization threshold"
condition_threshold {
filter = "resource.type = \"gce_instance\" AND metric.type = \"compute.googleapis.com/instance/cpu/utilization\""
comparison = "COMPARISON_GT"
threshold_value = var.cpu_threshold
duration = var.duration
aggregations {
alignment_period = var.alignment_period
per_series_aligner = "ALIGN_MEAN"
}
trigger { count = 1 }
}
}
notification_channels = google_monitoring_notification_channel.email[*].name
}
デフォルトの条件はこうです。CPU utilizationの5分平均が0.8を5分間超え、対象が1つ以上あればIncidentを開きます。
email[*].nameのSplat ExpressionはChannelが0件なら空List、1件ならResource名のListとなるため、同じAlert Policyで通知なし・ありの両方を扱えます。
設定と実行 #
Alert Policyだけの場合:
cd Basic-Examples/15-cloud-monitoring
cp terraform.tfvars.example terraform.tfvars
Email通知も作る場合:
notification_email = "you@example.com"
notification_channel_name = "tf-example-email"
terraform init
terraform fmt -check
terraform validate
terraform plan
terraform apply
Apply complete! Resources: 4 added, 0 changed, 0 destroyed.
Outputs:
alert_policy_name = "tf-example-gce-cpu-high"
notification_channel_name = "projects/YOUR_PROJECT_ID/notificationChannels/CHANNEL_ID"
GCP側で確認する #
gcloud monitoring policies list \
--project="$(terraform output -raw project_id)"
ConsoleではMonitoringのAlerting画面からConditionとNotification Channelを確認できます。Email ChannelはVerificationが必要になる場合があります。
後片付けと注意点 #
terraform destroy
Alert PolicyとChannelだけを削除します。監視設計ではThresholdだけでなく、No data時の扱い、通知先の冗長化、Incident対応手順、誤通知の抑制も検討します。
まとめ #
- Metric FilterとThresholdをTerraformで管理できる
- Alignment Periodと継続時間を変数化できる
countでEmail Channelを任意作成できる- Splat Expressionで0件または1件のChannelをPolicyへ渡せる
参考資料 #
次回 #
これでTerraform GCP Basic Examples 00〜15の基本サンプルを一通り確認できました。次は複数サービスを組み合わせるAdvanced Examplesへ進みます。