tags: master, kube-scheduler
本文档介绍部署高可用 kube-scheduler 集群的步骤。
该集群包含 3 个节点,启动后将通过竞争选举机制产生一个 leader 节点,其它节点为阻塞状态。当 leader 节点不可用后,剩余节点将再次进行选举产生新的 leader 节点,从而保证服务的可用性。
为保证通信安全,本文档先生成 x509 证书和私钥,kube-scheduler 在如下两种情况下使用该证书:
- 与 kube-apiserver 的安全端口通信;
- 在安全端口(https,10251) 输出 prometheus 格式的 metrics。
创建证书签名请求:
[admin@k8s-admin ~]$ cat >cert/kube-scheduler-csr.json <<EOF
{
"CN": "system:kube-scheduler",
"hosts": [
"127.0.0.1",
"192.168.99.101",
"192.168.99.102",
"192.168.99.103"
],
"key": {
"algo": "rsa",
"size": 2048
},
"names": [
{
"C": "CN",
"ST": "Jiangsu",
"L": "Suzhou",
"O": "system:kube-scheduler",
"OU": "wangxyd"
}
]
}
EOF
- hosts 列表包含所有 kube-scheduler 节点 IP;
- CN 为
system:kube-scheduler
、O 为system:kube-scheduler
,kubernetes 内置的 ClusterRoleBindingssystem:kube-scheduler
将赋予 kube-scheduler 工作所需的权限。
生成证书和私钥:
[admin@k8s-admin ~]$ cfssl gencert \
-ca=cert/ca.pem \
-ca-key=cert/ca-key.pem \
-config=cert/ca-config.json \
-profile=kubernetes \
cert/kube-scheduler-csr.json | cfssljson -bare cert/kube-scheduler
[admin@k8s-admin ~]$ ls cert/kube-scheduler*
cert/kube-scheduler.csr cert/kube-scheduler-key.pem
cert/kube-scheduler-csr.json cert/kube-scheduler.pem
kubeconfig 文件包含访问 apiserver 的所有信息,如 apiserver 地址、CA 证书和自身使用的证书:
[admin@k8s-admin ~]$ kubectl config set-cluster kubernetes \
--certificate-authority=cert/ca.pem \
--server=${KUBE_APISERVER} \
--embed-certs=true \
--kubeconfig=conf/kube-scheduler.kubeconfig
Cluster "kubernetes" set.
[admin@k8s-admin ~]$ kubectl config set-credentials system:kube-scheduler \
--client-certificate=cert/kube-scheduler.pem \
--client-key=cert/kube-scheduler-key.pem \
--embed-certs=true \
--kubeconfig=conf/kube-scheduler.kubeconfig
User "system:kube-scheduler" set.
[admin@k8s-admin ~]$ kubectl config set-context system:kube-scheduler \
--cluster=kubernetes \
--user=system:kube-scheduler \
--kubeconfig=conf/kube-scheduler.kubeconfig
Context "system:kube-scheduler" created.
[admin@k8s-admin ~]$ kubectl config use-context system:kube-scheduler \
--kubeconfig=conf/kube-scheduler.kubeconfig
Switched to context "system:kube-scheduler".
查看当前上下文的配置信息:
[admin@k8s-admin ~]$ kubectl config view \
--kubeconfig conf/kube-scheduler.kubeconfig
apiVersion: v1
clusters:
- cluster:
certificate-authority-data: DATA+OMITTED
server: https://192.168.99.100:8443
name: kubernetes
contexts:
- context:
cluster: kubernetes
user: system:kube-scheduler
name: system:kube-scheduler
current-context: system:kube-scheduler
kind: Config
preferences: {}
users:
- name: system:kube-scheduler
user:
client-certificate-data: REDACTED
client-key-data: REDACTED
分发 kubeconfig 到所有 master 节点:
[admin@k8s-admin ~]$ ansible k8s-masters -m copy -a "src=cert/kube-scheduler.kubeconfig dest=/etc/kubernetes/"
[admin@k8s-admin ~]$ cat >conf/kube-scheduler.service.j2 <<"EOF"
[Unit]
Description=Kubernetes Scheduler
Documentation=https://github.com/GoogleCloudPlatform/kubernetes
[Service]
WorkingDirectory={{ K8S_DIR }}/kube-scheduler
EnvironmentFile=/etc/kubernetes/kube-scheduler.conf
ExecStart=/usr/local/bin/kube-scheduler $KUBE_SCHEDULER_ARGS
Restart=on-failure
RestartSec=5
StartLimitInterval=0
[Install]
WantedBy=multi-user.target
EOF
[admin@k8s-admin ~]$ cat >conf/kube-scheduler.conf <<"EOF"
KUBE_SCHEDULER_ARGS="\
--port=0 \
--secure-port=10251 \
--bind-address=127.0.0.1 \
--kubeconfig=/etc/kubernetes/kube-scheduler.kubeconfig \
--leader-elect=true \
--kube-api-qps=100 \
--logtostderr=true \
--v=2"
EOF
[admin@k8s-admin ~]$ cat >conf/kube-scheduler.service.yaml <<"EOF"
- hosts: k8s-masters
remote_user: root
tasks:
- name: copy kube-scheduler.service
template:
src: kube-scheduler.service.j2
dest: /usr/lib/systemd/system/kube-scheduler.service
- name: copy kube-scheduler.conf
copy:
src: kube-scheduler.conf
dest: /etc/kubernetes/kube-scheduler.conf
- name: create WorkingDirectory
file:
path: '{{ K8S_DIR }}/kube-scheduler'
state: directory
- name: enable and start kube-scheduler.service
systemd:
name: kube-scheduler
state: restarted
enabled: yes
daemon_reload: yes
EOF
[admin@k8s-admin ~]$ ansible-playbook conf/kube-scheduler.service.yaml
[admin@k8s-admin ~]$ ansible k8s-masters -m shell -a "systemctl status kube-scheduler.service|grep -e Loaded -e Active"
- 必须创建工作目录。
- 确保状态为“active (running)”并且“enabled”。
参考测试 kube-controller-manager 集群的高可用,找一个或两个 master 节点,停掉 kube-scheduler 服务,看其它节点是否获取了 leader 权限(systemd 日志)。
查看当前的 leader 节点:
[admin@k8s-admin ~]$ kubectl get endpoints kube-scheduler --namespace=kube-system -o yaml
apiVersion: v1
kind: Endpoints
metadata:
annotations:
control-plane.alpha.kubernetes.io/leader: '{"holderIdentity":"k8s-node01_b234680e-3fdf-11e9-8410-525400e26008","leaseDurationSeconds":15,"acquireTime":"2019-03-06T07:16:02Z","renewTime":"2019-03-06T07:19:23Z","leaderTransitions":0}'
creationTimestamp: "2019-03-06T07:16:02Z"
name: kube-scheduler
namespace: kube-system
resourceVersion: "21838"
selfLink: /api/v1/namespaces/kube-system/endpoints/kube-scheduler
uid: b2cfb662-3fdf-11e9-98e2-525400bfa934
- 可见,当前的 leader 为 k8s-node01 节点。