# Alpha is failing on start

**URL:** <https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532>\
**Category:** Dgraph\
**Tags:** dgraph, status:accepted, kind:bug, area:kubernetes, ticket:created\
**Created:** [November 19, 2020, 10:31am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532 "2020-11-19T10:31:08Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 19, 2020, 10:31am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/1 "2020-11-19T10:31:08Z")

</div>

I set up a Dgraph cluster locally using minikube, 3x Alphas, 3x Zeros - everything was fine. Now I have scaled down all Alphas to 0

```auto
kubectl scale statefulset proj-graph-engine --replicas=0

```

then removed all the `pvc`s and `pv`s related to those Alphas and now when I’m scaling up the Alphas I get

```auto
...
I1119 10:26:06.414946 16 draft.go:1505] Calling IsPeer
E1119 10:26:06.415704 16 draft.go:1538] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
...

```

I’m using Dgraph v20.03.1

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 19, 2020, 11:47am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/2 "2020-11-19T11:47:01Z")

</div>

I have started a new cluster using Dgraph v20.03.6 and now after dropping PVCs and scaling Alphas up this shows up in the logs:

```auto
[pod/proj-graph-engine-zero-0/proj-graph-engine-zero] I1119 11:44:46.514916 17 zero.go:440] Connected: cluster_info_only:true
[pod/proj-graph-engine-1/proj-graph-engine] I1119 11:44:47.374876 15 draft.go:1543] Calling IsPeer
[pod/proj-graph-engine-1/proj-graph-engine] E1119 11:44:47.380827 15 draft.go:1576] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-zero-0/proj-graph-engine-zero] I1119 11:44:47.517125 17 zero.go:422] Got connection request: cluster_info_only:true
[pod/proj-graph-engine-0/proj-graph-engine] E1119 11:44:47.523360 18 draft.go:1576] Error while calling hasPeer: Unable to reach leader in group 1. Retrying...

```

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 19, 2020, 11:56am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/3 "2020-11-19T11:56:38Z")

</div>

And again, using Dgraph v20.07.2 and the same operation, scale down, remove PVC, scale up and then

```auto
[pod/proj-graph-engine-0/proj-graph-engine] I1119 11:55:31.892497 16 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-0/proj-graph-engine] E1119 11:55:31.894925 16 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-1/proj-graph-engine] I1119 11:55:31.990508 17 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-1/proj-graph-engine] E1119 11:55:31.996177 17 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-2/proj-graph-engine] I1119 11:55:32.865837 32 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-2/proj-graph-engine] E1119 11:55:32.879186 32 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-0/proj-graph-engine] I1119 11:55:32.896428 16 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-0/proj-graph-engine] E1119 11:55:32.904441 16 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-1/proj-graph-engine] I1119 11:55:32.996402 17 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-1/proj-graph-engine] E1119 11:55:33.000037 17 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-2/proj-graph-engine] I1119 11:55:33.880791 32 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-2/proj-graph-engine] E1119 11:55:33.886389 32 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-0/proj-graph-engine] I1119 11:55:33.904971 16 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-0/proj-graph-engine] E1119 11:55:33.912895 16 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...
[pod/proj-graph-engine-1/proj-graph-engine] I1119 11:55:34.004921 17 draft.go:1584] Calling IsPeer
[pod/proj-graph-engine-1/proj-graph-engine] E1119 11:55:34.010401 17 draft.go:1617] Error while calling hasPeer: error while joining cluster: rpc error: code = Unknown desc = No node has been set up yet. Retrying...

```

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 19, 2020, 12:01pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/4 "2020-11-19T12:01:19Z")

</div>

This means the official Upgrade Database procedure won’t work anymore  
[https://dgraph.io/docs/deploy/dgraph-administration/#upgrading-database](https://dgraph.io/docs/deploy/dgraph-administration/#upgrading-database)

---

<div class="post-metadata">

**Author:** ![MichelDiz](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/micheldiz/32/11873_2.png) [@MichelDiz](https://discuss.dgraph.io/u/MichelDiz)\
**Post date:** [November 23, 2020, 5:26pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/7 "2020-11-23T17:26:05Z")

</div>

I have accepted this to someone on the team to take a look.

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 23, 2020, 6:33pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/10 "2020-11-23T18:33:47Z")

</div>

I am looking at this right now, starting with `v20.03.6`.

---

<div class="post-metadata">

**Author:** ![dmai](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/dmai/32/1254_2.png) [@dmai](https://discuss.dgraph.io/u/dmai)\
**Post date:** [November 23, 2020, 6:41pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/11 "2020-11-23T18:41:39Z")

</div>

> [@lukaszlenart](#):
>
> then removed all the `pvc`s and `pv`s related to those Alphas and now when I’m scaling up the Alphas I get

@lukaszlenart Did you remove all the Alphas but keep the Zeros? If you’re looking to restart the cluster from scratch you’ll want to start from a clean slate (i.e., new data directories) for all Zeros and Alphas.

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 23, 2020, 7:00pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/12 "2020-11-23T19:00:58Z")

</div>

@lukaszlenart For this explicit process, using the dgraph helm chart, you could the following:

```bash
helm install pge --set image.tag=v20.03.6 dgraph/dgraph

## Scale Down Cluster and Delete Data + State
kubectl scale statefulset pge-dgraph-alpha --replicas=0
kubectl scale statefulset pge-dgraph-zero --replicas=0
kubectl delete pvc --selector release=pge

## Scale Up Cluster Starting with Zeros
kubectl scale statefulset pge-dgraph-zero --replicas=3

## Wait until 3 x healthy zero nodes
kubectl scale statefulset pge-dgraph-alpha --replicas=3

```

For the bulk loader, on an empty cluster, you would want to use an init container for bulk loader.

Generally, for immutable infrastructure patterns, It may be easier to just delete the statefulsets and recreate them from scratch again. With a helm chart used above, that process would be:

```bash
helm delete pge
kubectl delete pvc --selector release=pge

```

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 23, 2020, 7:02pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/13 "2020-11-23T19:02:37Z")

</div>

@dmai yes, I removed just Alphas, but then I removed Alphas and Zeros - problem persisted. The main issue is that, when you copied in all files from bulk loader to all the Alphas and then shut them down, they will start complaining in logs about `file already exists` and the cluster is broken.

@joaquin what do you mean by `initContainers`? copy data in?

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 23, 2020, 7:10pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/14 "2020-11-23T19:10:01Z")

</div>

@lukaszlenart Correct. On each of the Alphas, you’d have an initContainer, then do a spin-loop until you finish the bulk load and move directory created to `p`.

```auto
      command:
        - bash
        - "-c"
        - |
          trap "exit" SIGINT SIGTERM
          echo "Write to /dgraph/doneinit when ready."
          until [-f /dgraph/doneinit]; do sleep 2; done

```

Then `kubectl cp` file(s) into initContainer on alpha-0 pod (or from within the initContainer, `curl` it down), do the bulk load, `touch /dgraph/doneinit`. Do this same process for alpha-1, then alpha-2.

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 23, 2020, 7:13pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/15 "2020-11-23T19:13:31Z")

</div>

@joaquin thanks a lot, that should work!

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 23, 2020, 8:59pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/16 "2020-11-23T20:59:01Z")

</div>

@lukaszlenart As an example, I added initContainer automation in the current master of dgraph helm chart. If you wanted to use this, you could do the following.

### Get the Chart

```bash
git clone https://github.com/dgraph-io/charts.git

REL="pge"
helm install "REL" \
 --set image.tag=v20.03.6 \
 --set alpha.initContainers.init.enabled=true \
 ./charts/charts/dgraph/

```

### Copy Data to InitContainer

I used a sample dataset:

```bash
mkdir 1million && pushd 1million
PREFIX=https://github.com/dgraph-io/benchmarks/raw/master/data/
FILES=(1million.schema 1million.rdf.gz)

for FILE in ${FILES[*]}; do
  curl --silent --location --remote-name $PREFIX/$FILE
done

popd

```

Then I ran this process on alpha 0, 1, 2.

```bash
NUM=0
REL="pge"

kubectl cp ./1million/ $REL-dgraph-alpha-$NUM:/dgraph -c $REL-dgraph-alpha-init
kubectl exec -ti $REL-dgraph-alpha-$NUM -c $REL-dgraph-alpha-init -- bash

## inside initContainer
REL="pge"
dgraph bulk \
 --files /dgraph/1million/1million.rdf.gz \
 --schema /dgraph/1million/1million.schema \
 --zero $REL-dgraph-zero-0.$REL-dgraph-zero-headless.default.svc.cluster.local:5080

mv /dgraph/out/0/p /dgraph
touch doneinit

```

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 7:50am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/17 "2020-11-24T07:50:01Z")

</div>

Thanks a lot @joaquin! Just one question: can I run bulk loader on Alphas? I thought I need to do it on Zeros’ leader and the copy/paste “0” to all the Alphas.

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 7:58am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/18 "2020-11-24T07:58:24Z")

</div>

Hm… you run bulk loader in a dedicated init container which is just a Dgraph … interesting 🙂

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 8:14am UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/19 "2020-11-24T08:14:42Z")

</div>

One more question: is this Chart officially released?

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 12:33pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/20 "2020-11-24T12:33:30Z")

</div>

Tested and it works, osm!

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 24, 2020, 5:36pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/21 "2020-11-24T17:36:09Z")

</div>

Yes this chart is out there in the public domain.

- [dgraph 0.0.15 · helm/dgraph](https://artifacthub.io/packages/helm/dgraph/dgraph)

I didn’t announce the initContainer feature yet, as the interface will change from **`alpha.initContainers.generic.enabled`** to **`alpha.initContainers.init.enabled`**. I will also add further automation for specialized initContainers, such as _ **offline restore** _ and _ **bulkloader** _, but not sure if these two will make it to the next chart 0.0.13.

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 5:56pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/22 "2020-11-24T17:56:22Z")

</div>

Ach… I tried with `--set alpha.initContainers.init.enabled=true` and that’s why it didn’t work, thanks a lot!

---

<div class="post-metadata">

**Author:** ![joaquin](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/joaquin/32/2619_2.png) [@joaquin](https://discuss.dgraph.io/u/joaquin)\
**Post date:** [November 24, 2020, 6:00pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/23 "2020-11-24T18:00:47Z")

</div>

> [@lukaszlenart](#):
>
> Just one question: can I run bulk loader on Alphas? I thought I need to do it on Zeros’ leader and the copy/paste “0” to all the Alphas.

On this question, bulk loader can run anywhere, but it does need to connect to one of the Dgraph Zero nodes for the process to get timestamp generation. The zero leader is not needed, as members are equal partners and the leadership is elected (elected leader dependent on availability). This is part of the Raft consensus algorithm: [https://raft.github.io/](https://raft.github.io/).

The output (`./out`) that has the `p` directory (for 1 shard cluster) will need to be copied to each Dgraph Alpha node before it starts. Thus it should be possible to do it on one system, and copy the same `p` directory on each of the Dgraph Alpha nodes before they start.

I haven’t tried that exact process yet, as was following pattern how it would automated within Kubernetes (ala immutable infra style) for this.

---

<div class="post-metadata">

**Author:** ![lukaszlenart](https://yyz1.discourse-cdn.com/flex007/user_avatar/discuss.dgraph.io/lukaszlenart/32/2480_2.png) [@lukaszlenart](https://discuss.dgraph.io/u/lukaszlenart)\
**Post date:** [November 24, 2020, 7:02pm UTC](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532/24 "2020-11-24T19:02:56Z")

</div>

Thanks for the clarification. Does it mean I cannot run bulk loader on each Alpha and I must copy the `p` folder created during the first import to the rest of the Alphas?

[Next page](https://discuss.dgraph.io/t/alpha-is-failing-on-start/11532.md?page=2)
