# Us-central1 pangeo hub down?

**URL:** <https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591>\
**Category:** Uncategorized\
**Created:** [October 10, 2024, 12:47pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591 "2024-10-10T12:47:19Z")\
**Posts on this page:** 20\
**Page:** 2

<div class="post-metadata">

**Author:** ![ofk123](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ofk123/32/1853_2.png) [@ofk123](https://discourse.pangeo.io/u/ofk123)\
**Post date:** [October 14, 2024, 3:55pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/21 "2024-10-14T15:55:16Z")

</div>

It does not represent a solution to our blocked data-access, but since I am able to see the data-paths, it gives me confidence in the data still being there, which is what I was worried about the most.

- Hopefully there exists a way to regain access to folks data quickly!
- Getting the same hub back online, just for a week or two, would save our research project of a set-back. Probably other folks projects too? We just need to run one computation and store its output, in order to complete. I am hoping there exists some way to re-fund it. _crossing fingers_

Does someone in the Steering Council know if this is possible or not?  
And if not, we would appreciate some ideas into how we can move forward preferably without having to download and move data, if possible.

---

<div class="post-metadata">

**Author:** ![jmunroe](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/jmunroe/32/1632_2.png) [@jmunroe](https://discourse.pangeo.io/u/jmunroe)\
**Post date:** [October 14, 2024, 4:31pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/22 "2024-10-14T16:31:34Z")

</div>

As of this morning, there is a multipartner email thread making progress on getting access restored.

The current objective is to restore the access to the hub for at least a week so that those impacted have a chance to migrate their critical data.

Besides @ofk123 and @AndMei , if there are others who have been impacted by the loss of the Pangeo hub, please chime in on this thread.

---

<div class="post-metadata">

**Author:** ![lilmi](https://avatars.discourse-cdn.com/v4/letter/l/ec9cab/32.png) [@lilmi](https://discourse.pangeo.io/u/lilmi)\
**Post date:** [October 14, 2024, 5:26pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/23 "2024-10-14T17:26:39Z")

</div>

Hello and thank you for all your efforts on all of this! I wanted to add that I ( and several others in my group) have current work on Pangeo without any backup and we would be grateful for the chance to recover the data there!

---

<div class="post-metadata">

**Author:** ![ofk123](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ofk123/32/1853_2.png) [@ofk123](https://discourse.pangeo.io/u/ofk123)\
**Post date:** [October 14, 2024, 9:22pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/24 "2024-10-14T21:22:28Z")

</div>

Hi @jmunroe, our project group is discussing how to complete computations. Would it be possible to restore the hub with the same capacity it had last week, long enough for our computations to complete? If so we want to contribute to funding it.

Our parallelization is sort of tailor-written for the resources on Pangeo’s US Central, and our ~300GB of data is stored in an adjacent bucket. So a restoration would help us avoid migrating our data and finding a different computation-resource.

For the heaviest part of our remaining computation, we scale to ~1500 workers, each with 1 CPU and 7 GB RAM (the initial US Central configuration). I have not calculated how long, but likely ~5-10 hours in total. So a week should potentially be sufficient for us to complete our work.

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 15, 2024, 5:44am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/25 "2024-10-15T05:44:53Z")

</div>

Thanks massively for this, I’ll be ready to migrate my data when you let us know. Thanks again to all for their efforts.

---

<div class="post-metadata">

**Author:** ![rabernat](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/rabernat/32/22_2.png) [@rabernat](https://discourse.pangeo.io/u/rabernat)\
**Post date:** [October 15, 2024, 6:43pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/26 "2024-10-15T18:43:14Z")

</div>

Here’s an update–we have identified an interim funding source, and CUIT is in the process of reactivating the accounts.

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 15, 2024, 9:05pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/27 "2024-10-15T21:05:20Z")

</div>

This is terrific news. Thanks for being so proactive with this Ryan!

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 16, 2024, 9:14am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/28 "2024-10-16T09:14:53Z")

</div>

The hub is now back up!

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 16, 2024, 9:34am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/29 "2024-10-16T09:34:08Z")

</div>

Almost. I can now connect, but upon starting the smallest server available it eventually times out with the error:

“2024-10-16T09:30:36Z [Warning] MountVolume.SetUp failed for volume “prod-home-nfs” : mount failed: exit status 1 Mounting command: /home/kubernetes/containerized\_mounter/mounter Mounting arguments: mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/56cb0143-d1d4-4038-a4ac-57046a02c03d/volumes/kubernetes.io~nfs/prod-home-nfs Output: Mount failed: mount failed: exit status 32 Mounting command: chroot Mounting arguments: [/home/kubernetes/containerized\_mounter/rootfs mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/56cb0143-d1d4-4038-a4ac-57046a02c03d/volumes/kubernetes.io~nfs/prod-home-nfs] Output: mount.nfs: Connection timed out”

I’ll keep trying, but I’m not sure it is fully working yet!

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 16, 2024, 9:35am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/30 "2024-10-16T09:35:41Z")

</div>

I am just deploying a change to the homepage to alert people to the timeline so I can look into this. I was going off our monitoring system that reported the URL resolving again.

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 16, 2024, 9:42am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/31 "2024-10-16T09:42:43Z")

</div>

Full error log here:

" Event log

Server requested

2024-10-16T09:26:19Z [Warning] 0/3 nodes are available: 1 node(s) had untolerated taint {ToBeDeletedByClusterAutoscaler: 1729070778}, 2 node(s) didn’t match Pod’s node affinity/selector. preemption: 0/3 nodes are available: 3 Preemption is not helpful for scheduling.

2024-10-16T09:26:27Z [Normal] pod triggered scale-up: [{[https://www.googleapis.com/compute/v1/projects/pangeo-integration-te-3eea/zones/us-central1-b/instanceGroups/gke-pangeo-hubs-cluster-nb-small-c97e04c1-grp](https://www.googleapis.com/compute/v1/projects/pangeo-integration-te-3eea/zones/us-central1-b/instanceGroups/gke-pangeo-hubs-cluster-nb-small-c97e04c1-grp) 0-\>1 (max: 100)}]

2024-10-16T09:27:34Z [Normal] Successfully assigned prod/jupyter-recalculate to gke-pangeo-hubs-cluster-nb-small-c97e04c1-vhww

2024-10-16T09:27:34Z [Normal] Successfully assigned prod/jupyter-recalculate to gke-pangeo-hubs-cluster-nb-small-c97e04c1-vhww

2024-10-16T09:30:36Z [Warning] MountVolume.SetUp failed for volume “prod-home-nfs” : mount failed: exit status 1 Mounting command: /home/kubernetes/containerized\_mounter/mounter Mounting arguments: mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/56cb0143-d1d4-4038-a4ac-57046a02c03d/volumes/kubernetes.io~nfs/prod-home-nfs Output: Mount failed: mount failed: exit status 32 Mounting command: chroot Mounting arguments: [/home/kubernetes/containerized\_mounter/rootfs mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/56cb0143-d1d4-4038-a4ac-57046a02c03d/volumes/kubernetes.io~nfs/prod-home-nfs] Output: mount.nfs: Connection timed out

Spawn failed: Timeout"

Thanks for being proactive on this!

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 16, 2024, 9:45am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/32 "2024-10-16T09:45:59Z")

</div>

_gulp_

 ![Screenshot 2024-10-16 at 10.44.31](https://canada1.discourse-cdn.com/flex030/uploads/pangeo/original/2X/a/aff76fb9605626633a3e0a4c6b8ceaa2d2ff0649.png)

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 16, 2024, 9:52am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/33 "2024-10-16T09:52:21Z")

</div>

Does this mean the server is up but the data doesn’t exist or is not linked?

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 16, 2024, 9:53am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/34 "2024-10-16T09:53:27Z")

</div>

I think it’s a networking issue that I don’t know how to debug. I tried the browser debugging stuff, and it didn’t help. So I suspect something has changed to affect the network.

ETA: I’ve emailed Columbia IT again.

---

<div class="post-metadata">

**Author:** ![ofk123](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ofk123/32/1853_2.png) [@ofk123](https://discourse.pangeo.io/u/ofk123)\
**Post date:** [October 16, 2024, 7:53pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/35 "2024-10-16T19:53:07Z")

</div>

Just adding that the connection-timeout also occurs on my end.

> **Event log**
>
> ```auto
> Server requested
> 2024-10-16T19:37:44Z [Warning] 0/2 nodes are available: 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling.
> 2024-10-16T19:37:51Z [Normal] pod triggered scale-up: [{https://www.googleapis.com/compute/v1/projects/pangeo-integration-te-3eea/zones/us-central1-b/instanceGroups/gke-pangeo-hubs-cluster-nb-medium-d51fa3b8-grp 0->1 (max: 100)}]
> 2024-10-16T19:38:39Z [Normal] Successfully assigned prod/jupyter-ofk123 to gke-pangeo-hubs-cluster-nb-medium-d51fa3b8-vlq6
> 2024-10-16T19:38:41Z [Normal] Cancelling deletion of Pod prod/jupyter-ofk123
> 2024-10-16T19:44:45Z [Warning] MountVolume.SetUp failed for volume "prod-home-nfs" : mount failed: exit status 1 Mounting command: /home/kubernetes/containerized_mounter/mounter Mounting arguments: mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/6cf22ac1-1cda-4927-aabf-2350da6fc999/volumes/kubernetes.io~nfs/prod-home-nfs Output: Mount failed: mount failed: exit status 32 Mounting command: chroot Mounting arguments: [/home/kubernetes/containerized_mounter/rootfs mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/6cf22ac1-1cda-4927-aabf-2350da6fc999/volumes/kubernetes.io~nfs/prod-home-nfs] Output: mount.nfs: Connection timed out
> Spawn failed: pod prod/jupyter-ofk123 did not start in 600 seconds!
> 
> ```

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 17, 2024, 7:08am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/36 "2024-10-17T07:08:17Z")

</div>

Does the previous method you tried for accessing the buckets work now that the billing account has been reactivated? (I realise this is not a solution for accessing anything that was in your home directory.)

---

<div class="post-metadata">

**Author:** ![AndMei](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@AndMei](https://discourse.pangeo.io/u/AndMei)\
**Post date:** [October 17, 2024, 11:56am UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/37 "2024-10-17T11:56:41Z")

</div>

Great, thanks for doing so. Fingers crossed for a quick response!

---

<div class="post-metadata">

**Author:** ![ofk123](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ofk123/32/1853_2.png) [@ofk123](https://discourse.pangeo.io/u/ofk123)\
**Post date:** [October 17, 2024, 2:09pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/38 "2024-10-17T14:09:36Z")

</div>

Thanks, yes, now I am able to access data from EOSC. There is no longer an `OSError` .

---

<div class="post-metadata">

**Author:** ![ofk123](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/ofk123/32/1853_2.png) [@ofk123](https://discourse.pangeo.io/u/ofk123)\
**Post date:** [October 17, 2024, 2:37pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/39 "2024-10-17T14:37:29Z")

</div>

But the same timeout occurs on the US Central hub today.

> **Event log**
>
> ```auto
> Server requested
> 2024-10-17T14:08:19Z [Warning] 0/2 nodes are available: 2 node(s) didn't match Pod's node affinity/selector. preemption: 0/2 nodes are available: 2 Preemption is not helpful for scheduling.
> 2024-10-17T14:08:21Z [Normal] pod triggered scale-up: [{https://www.googleapis.com/compute/v1/projects/pangeo-integration-te-3eea/zones/us-central1-b/instanceGroups/gke-pangeo-hubs-cluster-nb-medium-d51fa3b8-grp 0->1 (max: 100)}]
> 2024-10-17T14:09:09Z [Normal] Successfully assigned prod/jupyter-ofk123 to gke-pangeo-hubs-cluster-nb-medium-d51fa3b8-87gb
> 2024-10-17T14:09:12Z [Normal] Cancelling deletion of Pod prod/jupyter-ofk123
> 2024-10-17T14:12:14Z [Warning] MountVolume.SetUp failed for volume "prod-home-nfs" : mount failed: exit status 1 Mounting command: /home/kubernetes/containerized_mounter/mounter Mounting arguments: mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/bdf5296f-2a93-410e-98bc-9d06b2865333/volumes/kubernetes.io~nfs/prod-home-nfs Output: Mount failed: mount failed: exit status 32 Mounting command: chroot Mounting arguments: [/home/kubernetes/containerized_mounter/rootfs mount -t nfs -o noatime,soft 10.229.44.234:/homes/prod /var/lib/kubelet/pods/bdf5296f-2a93-410e-98bc-9d06b2865333/volumes/kubernetes.io~nfs/prod-home-nfs] Output: mount.nfs: Connection timed out
> Spawn failed: Timeout
> 
> ```

Hopefully you get some answers from the IT-department soon. Thanks again.

---

<div class="post-metadata">

**Author:** ![sgibson91](https://yyz2.discourse-cdn.com/flex030/user_avatar/discourse.pangeo.io/sgibson91/32/1051_2.png) [@sgibson91](https://discourse.pangeo.io/u/sgibson91)\
**Post date:** [October 17, 2024, 4:14pm UTC](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591/40 "2024-10-17T16:14:06Z")

</div>

Yes, I’m still waiting on information regarding the NFS server for home directories

[Previous page](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591.md?page=1)

[Next page](https://discourse.pangeo.io/t/us-central1-pangeo-hub-down/4591.md?page=3)
