概要 #
GKEはコンソールから作ると、決めることが多いサービスです。Node PoolやWorkload Identityなど、何が有効か把握しづらくなります。またNodeが起動していても、そのNode上でPodが実際に動くかどうかは、確認しなければわからないままです。
Terraformならこれらを一式コードに残せます。まとめて作り、まとめて消せる構成です。ここでは最小構成のZonal GKE Standard ClusterとSpot Node Pool(1ノード)を作成します。実際にPodを動かして確認します。
Terraformの基本操作と、kubectlでの基本的なリソース操作(get、run、delete)は前提としています。
Ingress、Cloud NAT、実際のWorkloadアプリのデプロイ、Autopilotへの変更は扱いません。学習・短時間検証向けの最小構成です。
検証すること #
- VPC-nativeのGKE Standard Clusterを作成できること
- Default Node Poolを分離し、Spot Node Poolを管理できること
- gcloudとkubectlでClusterとNodeの状態を確認できること
- Spot Node Pool上にPodを実際にスケジュールできること(Nodeの起動確認とは別)
- terraform destroyで検証Resourceを削除できること
今回の構成 #
まず、何がどこに置かれるかです。Secondary RangeはSubnetの属性です。GKE ClusterとSpot Node Poolは、そのSubnet上に作られます。
flowchart TB
subgraph GCP["Google Cloud"]
subgraph Project["Project"]
subgraph Region["asia-northeast1"]
subgraph VPC["VPC"]
subgraph Subnet["Subnet: 10.40.0.0/24"]
PodsRange["Pods secondary range
10.41.0.0/16"]
SvcRange["Services secondary range
10.42.0.0/20"]
subgraph Cluster["GKE Cluster: tf-example-gke
zone asia-northeast1-a / REGULAR Channel"]
CP["Control Plane"]
subgraph Pool["Spot Node Pool: spot-pool"]
Node["Node
e2-medium"]
end
end
end
end
end
end
end
CP -.->|"管理"| Pool
PodとServiceはSubnetのSecondary Rangeから払い出されます。Primaryと重複しないCIDR設計が必要です。
実際にPodを動かすところまでの流れは、次のようになります。
sequenceDiagram
actor User as 手元の端末
participant API as Control Plane
participant Sched as Scheduler
participant Node as Node(kubelet)
User->>API: gcloud container clusters get-credentials
User->>API: kubectl run test-pod --image=nginx
API->>Sched: Pod作成をScheduler へ通知
Sched->>Node: Spot Node PoolのNodeへ割り当て
Node-->>API: Pod Running を報告
API-->>User: kubectl get pod → Running
前提環境 #
- Google Cloud CLI、Terraform、kubectl
- Billingが有効な検証用Google Cloud Project
- GKE、Compute Engine、IAMを操作できるGoogleアカウント
- 検証時のバージョン:Terraform 1.14.3、google 7.43.0
使用するTerraformコード #
VPC-native用Secondary Range #
PodとServiceへSubnetのSecondary Rangeを割り当てます。
resource "google_compute_subnetwork" "primary" {
ip_cidr_range = var.subnet_cidr
region = var.region
network = google_compute_network.vpc.id
secondary_ip_range {
range_name = "pods"
ip_cidr_range = var.pods_cidr
}
secondary_ip_range {
range_name = "services"
ip_cidr_range = var.services_cidr
}
}
Primary、Pods、ServicesのCIDRが重複しないよう設計します。
GKE Cluster #
resource "google_container_cluster" "primary" {
name = var.cluster_name
location = var.zone
network = google_compute_network.vpc.name
subnetwork = google_compute_subnetwork.primary.name
remove_default_node_pool = true
initial_node_count = 1
deletion_protection = false
release_channel { channel = "REGULAR" }
ip_allocation_policy {
cluster_secondary_range_name = "pods"
services_secondary_range_name = "services"
}
workload_identity_config {
workload_pool = "${var.project_id}.svc.id.goog"
}
}
Terraform Providerの要件により、一時的なDefault Node Poolを作成します。すぐ削除し、別ResourceのSpot Node Poolへ置き換える構成です。Workload Identity Federation for GKEも有効にしています。
Spot Node Pool #
resource "google_container_node_pool" "spot" {
name = "spot-pool"
cluster = google_container_cluster.primary.name
node_count = var.node_count
node_config {
machine_type = var.machine_type
spot = true
disk_size_gb = 30
disk_type = "pd-balanced"
workload_metadata_config { mode = "GKE_METADATA" }
}
management {
auto_repair = true
auto_upgrade = true
}
}
Spot Nodeは回収される可能性があります。可用性が必要なWorkloadには、On-demand Node Poolや複数Nodeを検討します。
設定と実行 #
cd Basic-Examples/07-gke
cp terraform.tfvars.example terraform.tfvars
terraform init
terraform fmt -check
terraform validate
terraform plan
terraform apply
VPC、Subnet、Cluster、Spot Node Poolなど6つが作成されます。Node Poolの作成には1分強かかりました。
Apply complete! Resources: 6 added, 0 changed, 0 destroyed.
Outputs:
cluster_location = "asia-northeast1-a"
cluster_name = "tf-example-gke"
get_credentials_example = "gcloud container clusters get-credentials tf-example-gke --zone=asia-northeast1-a --project=YOUR_PROJECT_ID"
network_name = "tf-example-gke-vpc"
GCPとKubernetes側で確認する #
gcloud container clusters describe \
"$(terraform output -raw cluster_name)" \
--zone="$(terraform output -raw cluster_location)" \
--project=YOUR_PROJECT_ID
Cluster / Network / Secondary Rangeの対応を抜粋します。
name: tf-example-gke
location: asia-northeast1-a
network: tf-example-gke-vpc
networkConfig:
subnetwork: projects/YOUR_PROJECT_ID/regions/asia-northeast1/subnetworks/tf-example-gke-subnet
ipAllocationPolicy:
clusterSecondaryRangeName: pods
servicesSecondaryRangeName: services
nodePools:
- name: spot-pool
config:
machineType: e2-medium
spot: true
Credentialを取得してNodeを確認します。
gcloud container clusters get-credentials \
"$(terraform output -raw cluster_name)" \
--zone="$(terraform output -raw cluster_location)" \
--project=YOUR_PROJECT_ID
kubectl get nodes -o wide
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
gke-tf-example-gke-spot-pool-84c7fe12-fnzh Ready <none> 33s v1.35.7-gke.1027000 10.40.0.4 34.146.254.71 Container-Optimized OS from Google 6.12.94+ containerd://2.1.9
Podを実際に動かして確認する #
NodeがReadyであることは、そのNode上でPodが実際に動くことを保証しません。イメージのpullやスケジューリングが通るかは、Podを1つ動かして初めて分かります。
kubectl run test-pod --image=nginx --restart=Never
kubectl wait --for=condition=Ready pod/test-pod --timeout=120s
kubectl get pod test-pod -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
test-pod 1/1 Running 0 18s 10.41.0.13 gke-tf-example-gke-spot-pool-84c7fe12-fnzh <none> <none>
STATUSがRunningになり、IPにはPods Secondary Range(10.41.0.0/16)内のアドレスが割り当てられています。確認できたら削除します。
kubectl delete pod test-pod
後片付け #
terraform destroy
Destroy complete! Resources: 6 destroyed.
GKE Cluster、Node、Diskは課金対象です。検証後は必ずdestroyし、gcloud container clusters listでClusterが消えていることも確認します。
まとめ #
- VPC-native用にPod・ServiceのSecondary Rangeを設定できる
- Default Node Poolを分離して管理できる
- Spot Node Poolで学習用構成を小さくできる
- Workload IdentityとRelease Channelを有効化できる
- Nodeの起動確認とは別に、Podが実際にスケジュールされることを確認できる
参考資料 #
次回 #
次はPrivate Cloud DNS ZoneとA RecordをTerraformで作成します。