<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel>
  <title>Mohamed Elsakhawy, PhD</title>
  <link>https://mohamede.com/</link>
  <description>Cloud computing research, systems engineering, and talks</description>
  <language>en</language>
  <atom:link href="https://mohamede.com/feed.xml" rel="self" type="application/rss+xml"/>
  <item>
    <title>My talk at SCaLE 22x</title>
    <link>https://mohamede.com/posts/my-talk-at-scalex-22/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/my-talk-at-scalex-22/</guid>
    <pubDate>Sun, 09 Mar 2025 05:06:19 +0000</pubDate>
    <description>Read More</description>
    <content:encoded><![CDATA[<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<div class="embed-container"><div class="video"><iframe src="https://www.youtube-nocookie.com/embed/wFW7V5ahYdI?feature=oembed" title="Video" loading="lazy" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></div></div>
</div></figure>]]></content:encoded>
  </item>
  <item>
    <title>SCALE 22x</title>
    <link>https://mohamede.com/posts/scale-22x/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/scale-22x/</guid>
    <pubDate>Fri, 13 Dec 2024 09:40:33 +0000</pubDate>
    <description>I am pleased to announce that I will be speaking at SCALE 22x between March 6 and 9th , 2025 at the Pasadena Convention Center in Pasadena, CA. More to follow ! SCALE 22x Read More</description>
    <content:encoded><![CDATA[<p class="wp-block-paragraph">I am pleased to announce that I will be speaking at 
 <a href="https://www.socallinuxexpo.org/scale/22x/presentations/deep-dive-linux-networking" target="_blank" rel="noopener noreferrer"><span class="hl-blue">SCALE 22x </span></a>  between March 6 and 9th , 2025 at the Pasadena Convention Center in Pasadena, CA. More to follow !</p>



 <a href="https://www.socallinuxexpo.org/scale/22x/presentations/deep-dive-linux-networking" target="_blank" rel="noopener noreferrer"><span class="hl-blue">SCALE 22x </span></a>  







<figure class="wp-block-image size-full"><a href="https://www.socallinuxexpo.org/scale/22x/presentations/deep-dive-linux-networking" target="_blank" rel="noopener noreferrer"><picture><source type="image/webp" srcset="https://mohamede.com/assets/uploads/2024/12/image-480.webp 480w, assets/uploads/2024/12/image-760.webp 760w, assets/uploads/2024/12/image-1520.webp 1520w" sizes="(max-width: 792px) 100vw, 760px"><img src="https://mohamede.com/assets/uploads/2024/12/image.png" alt="" width="3500" height="1387" loading="lazy" decoding="async"></picture></a></figure>]]></content:encoded>
  </item>
  <item>
    <title>My talk at OpenInfra Days North America</title>
    <link>https://mohamede.com/posts/my-talk-at-openinfra-days-north-america/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/my-talk-at-openinfra-days-north-america/</guid>
    <pubDate>Sun, 20 Oct 2024 14:26:11 +0000</pubDate>
    <description>I am pleased to speak at OpenInfra Days North America at Indiana University. The link for the talk Read More</description>
    <content:encoded><![CDATA[<p class="wp-block-paragraph">I am pleased to speak at OpenInfra Days North America at Indiana University.  </p>



<a href="https://oidiu2024.sched.com/event/1fH6e/back-to-the-fundamentals-are-you-actually-monitoring-your-cloud" target="_blank" rel="noopener noreferrer"><span class="hl-blue">The link for the talk</a>]]></content:encoded>
  </item>
  <item>
    <title>Paper Accepted in ATC USENIX</title>
    <link>https://mohamede.com/posts/paper-accepted-in-atc-usenix/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/paper-accepted-in-atc-usenix/</guid>
    <pubDate>Sat, 08 Jul 2023 15:27:00 +0000</pubDate>
    <description>I am pleased to announce that our paper was accepted at the ATC USENIX’23 conference in Boston “UnFaaSener: Latency and Cost Aware Offloading of Functions from Serverless Platforms“ Read More</description>
    <content:encoded><![CDATA[<p>I am pleased to announce that our <strong><a title="" href="https://search.app?link=https%3A%2F%2Fwww.usenix.org%2Fconference%2Fatc23%2Fpresentation%2Fsadeghian&amp;utm_campaign=aga&amp;utm_source=agsadl1%2Csh%2Fx%2Fgs%2Fm2%2F4"><span class="hl-blue">paper</span></a> </strong>was accepted at the ATC USENIX&#8217;23 conference in Boston</p>



<p class="wp-block-paragraph"><strong><a title="" href="https://search.app/?link=https%3A%2F%2Fwww.usenix.org%2Fconference%2Fatc23%2Fpresentation%2Fsadeghian&amp;utm_campaign=aga&amp;utm_source=agsadl1%2Csh%2Fx%2Fgs%2Fm2%2F4">&#8220;<em><span class="hl-blue">UnFaaSener: Latency and Cost Aware Offloading of Functions from Serverless Platforms</em>&#8220;</a> </strong></p>]]></content:encoded>
  </item>
  <item>
    <title>Paper accepted at WoSC ‘7</title>
    <link>https://mohamede.com/posts/paper-accepted-at-wosc-7/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/paper-accepted-at-wosc-7/</guid>
    <pubDate>Mon, 24 Jan 2022 09:44:00 +0000</pubDate>
    <description>Read More</description>
    <content:encoded><![CDATA[<figure class="wp-block-embed"><div class="wp-block-embed__wrapper">
https://www.serverlesscomputing.org/wosc7/papers/p6
</div></figure>



<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<div class="embed-container"><div class="video"><iframe src="https://www.youtube-nocookie.com/embed/Us_1-HwfLxA?start=1&amp;feature=oembed" title="Video" loading="lazy" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></div></div>
</div></figure>]]></content:encoded>
  </item>
  <item>
    <title>Open Infrastructure Summit , Berlin 2020</title>
    <link>https://mohamede.com/posts/open-infrastructure-summit-berlin-2020/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/open-infrastructure-summit-berlin-2020/</guid>
    <pubDate>Thu, 16 Jul 2020 07:25:44 +0000</pubDate>
    <description>I’m pleased to serve on the programming committee of the “Getting Started” track in the upcoming Open Infra summit in Berlin. If you plan to submit a talk and have questions, looking for advice or just want to have a…</description>
    <content:encoded><![CDATA[<p>I&#8217;m pleased to serve on the programming committee of the <strong> &#8220;Getting Started&#8221; </strong> track in the upcoming <a href="https://www.openstack.org/summit/2020/" target="_blank" rel="noopener noreferrer">Open Infra summit in Berlin</a>. If you plan to submit a talk and have questions, looking for advice or just want to have a chat on your proposal. I will be happy to help you with crafting the proposal.</p>
<p>My office hours are:</p>
<p>11 &#8211; 12:30 UTC, Thursdays on <a href="https://webchat.freenode.net/" target="_blank" rel="noopener noreferrer">Freenode</a> IRC, #open-infra-summit-cfp</p>
<p>Good Luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>Openstack UC and TC to Unite!</title>
    <link>https://mohamede.com/posts/uc-and-tc-to-unite/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/uc-and-tc-to-unite/</guid>
    <pubDate>Sun, 28 Jun 2020 09:36:44 +0000</pubDate>
    <description>This coming month marks the last running session of the Openstack User Committee. Over the years, the OpenStack community has grown with many operators being directly involved in the development lifecycle. In efforts to…</description>
    <content:encoded><![CDATA[<p>This coming month marks the last running session of the Openstack User Committee. Over the years, the OpenStack community has grown with many operators being directly involved in the development lifecycle. In efforts to cope with such change, we needed to adjust the governance model to remove the barriers between the governing bodies, and in turn enabling more involvement from operators in the various projects. Thus, starting on August 1st,  <a href="http://lists.openstack.org/pipermail/openstack-discuss/2020-June/015604.html" target="_blank" rel="noopener noreferrer">the UC will unite with TC to be a single governance body under TC.</a></p>
<p>I am honored to have been part of the UC and to have served as the chair in its last round. I&#8217;d like to thank all current and past UC members for their efforts that together have supported the Openstack user community over the past years. I&#8217;d like also to thank the user community for entrusting the UC and supporting its mission in serving and representing all Openstack users. I&#8217;m confident that the united body will do a great job and continue to provide a strong representation of the user community and serve its needs.</p>
<p>So long UC, and thanks for all the fish !</p>]]></content:encoded>
  </item>
  <item>
    <title>Cephadm: Good bye ceph-deploy</title>
    <link>https://mohamede.com/posts/cephadm-so-long-ceph-deploy/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/cephadm-so-long-ceph-deploy/</guid>
    <pubDate>Tue, 23 Jun 2020 18:15:07 +0000</pubDate>
    <description>As you probably may know, ceph-deploy, the beloved deployment utility for CEPH, is no longer maintained. Cephadm is the new tool/package to deploy CEPH clusters. CERN has a pretty good intro PDF to it. Cephadm includes…</description>
    <content:encoded><![CDATA[<p>As you probably may know, ceph-deploy, the beloved deployment utility for CEPH, is <a href="https://ceph.readthedocs.io/en/latest/install/" target="_blank" rel="noopener noreferrer">no longer maintained</a>. Cephadm is the new tool/package to deploy CEPH clusters.</p>
<p>CERN has a pretty good intro <a href="https://indico.cern.ch/event/898515/attachments/2011108/3360279/Ceph_Octopus__Orch.pdf" target="_blank" rel="noopener noreferrer">PDF</a> to it.  Cephadm includes many nice features including the ability to adopt running CEPH clusters.</p>
<p>Two quick notes that&#8217;ll save you some time</p>
<ul>
<li>While adding hosts using the</li>
</ul>
<div class="codeblock"><pre><code>ceph orch host add hostname</code></pre></div>
<p>You need to specify the IP of the host as follows</p>
<div class="codeblock"><pre><code>ceph orch host add hostname IP.IP.IP.IP</code></pre></div>
<p>if you get the following error, despite injecting the ssh keys correctly</p>
<div class="codeblock"><pre><code>Error ENOENT: Failed to connect to hostname (hostname).  Check that the host is reachable and accepts connections using the cephadm SSH key
you may want to run: 
&gt; ssh -F =(ceph cephadm get-ssh-config) -i =(ceph config-key get mgr/cephadm/ssh_identity_key) root@hostname</code></pre></div>
<ul>
<li>When adding OSDs</li>
</ul>
<p>If you are deploying a cluster with a &#8220;relatively&#8221; moderate number of OSDs per host, you may run into the following error scenario while using:</p>
<div class="codeblock"><pre><code> ceph orch apply osd --all-available-devices</code></pre></div>
<p>The command basically adds the available hdds/ssds to be part of your cluster. Under the hood, this is done by running a docker container that&#8217;s in charge of that OSD. Basically the following command is run</p>
<div class="codeblock"><pre><code>/bin/bash /var/lib/ceph/{FSID}/osd.{NUM}/unit.run</code></pre></div>
<p>It does that for every available OSD in your hosts. You may find that some of the OSDs don&#8217;t start and are stuck in error start despite your efforts to use</p>
<div class="codeblock"><pre><code>ceph orch daemon restart osd.xx</code></pre></div>
<p>If you dig deeper (by executing the docker shell directly or looking into the logs) , you will find the following self-explanatory error</p>
<div class="codeblock"><pre><code> /var/lib/ceph/osd/ceph-xx/block) _aio_start io_setup(2) failed with EAGAIN; try increasing /proc/sys/fs/aio-max-nr</code></pre></div>
<p>The solution is simply to set the asynchronous non-blocking IO into a higher value using</p>
<div class="codeblock"><pre><code>sudo sysctl -w fs.aio-max-nr=1048576</code></pre></div>
<p>If that solves your issue, apply it to sysctl.conf to persist</p>
<p>Happy cephadmining 🙂</p>]]></content:encoded>
  </item>
  <item>
    <title>My talk at Stackconf</title>
    <link>https://mohamede.com/posts/my-talk-at-stackconf/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/my-talk-at-stackconf/</guid>
    <pubDate>Thu, 18 Jun 2020 12:12:18 +0000</pubDate>
    <description>Pleased to speak at Stackconf.eu this year Direct Youtube link Read More</description>
    <content:encoded><![CDATA[<p>Pleased to speak at Stackconf.eu this year</p>
<p><a href="https://youtu.be/t40oTtUse84" target="_blank" rel="noopener noreferrer"><span class="hl-blue">Direct Youtube link</span></a></p>


<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<div class="embed-container"><div class="video"><iframe src="https://www.youtube-nocookie.com/embed/t40oTtUse84?feature=oembed" title="Video" loading="lazy" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></div></div>
</div></figure>]]></content:encoded>
  </item>
  <item>
    <title>Running pods on master nodes in RKE</title>
    <link>https://mohamede.com/posts/running-pods-on-master-nodes-in-rke/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/running-pods-on-master-nodes-in-rke/</guid>
    <pubDate>Sun, 29 Mar 2020 06:28:48 +0000</pubDate>
    <description>You may run into situations where you need to run pods on your K8s master node. If you’r using RKE, you need to taint two labels on the controller/master node. You can obtain the current labels using kubectl describe…</description>
    <content:encoded><![CDATA[<p>You may run into situations where you need to run pods on your K8s master node. If you&#8217;r using RKE, you need to taint two labels on the controller/master node.</p>
<p>You can obtain the current labels using</p>
<div class="codeblock"><pre><code>kubectl describe node NODENAME</code></pre></div>
<p>and look under the labels section</p>
<div class="codeblock"><pre><code>Labels: beta.kubernetes.io/arch=amd64
beta.kubernetes.io/os=linux
kubernetes.io/arch=amd64
kubernetes.io/hostname=lab-1
kubernetes.io/os=linux
node-role.kubernetes.io/controlplane=true
node-role.kubernetes.io/etcd=true</code></pre></div>
<p>you then need the following two commands</p>
<div class="codeblock"><pre><code>kubectl taint nodes node-1 node-role.kubernetes.io/etcd-
kubectl taint nodes node-1 node-role.kubernetes.io/controlplane-</code></pre></div>
<p>Happy Kuberneting!</p>]]></content:encoded>
  </item>
  <item>
    <title>Speaking at Stackconf !</title>
    <link>https://mohamede.com/posts/presenting-at-stackconf/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/presenting-at-stackconf/</guid>
    <pubDate>Fri, 28 Feb 2020 17:10:30 +0000</pubDate>
    <description>I’m happy to announce that I will be speaking at the upcoming Stackconf held in Berlin in June 2020. Stay tuned ! Read More</description>
    <content:encoded><![CDATA[<p>I&#8217;m happy to announce that I will be speaking at the upcoming Stackconf held in Berlin in June 2020. <a href="https://stackconf.eu/talks/opensource-in-advanced-research-computing-how-canada-did-it/" target="_blank" rel="noopener noreferrer">Stay tuned ! </a></p>]]></content:encoded>
  </item>
  <item>
    <title>Our paper in IEEE Smart Cloud 2019</title>
    <link>https://mohamede.com/posts/our-paper-in-ieee-smart-cloud-2019/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/our-paper-in-ieee-smart-cloud-2019/</guid>
    <pubDate>Wed, 18 Dec 2019 20:52:49 +0000</pubDate>
    <description>Pleased to present our latest work in the IEEE Smart Cloud Conference in Tokyo https://ieeexplore.ieee.org/document/9091413 Read More</description>
    <content:encoded><![CDATA[<p>Pleased to present our latest work in the IEEE Smart Cloud Conference in Tokyo</p>
<p><a href="https://ieeexplore.ieee.org/document/9091413" target="_blank" rel="noopener noreferrer">https://ieeexplore.ieee.org/document/9091413</a></p>]]></content:encoded>
  </item>
  <item>
    <title>Apache Openwhisk with Kubespray</title>
    <link>https://mohamede.com/posts/apache-openwhisk-with-kubespray/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/apache-openwhisk-with-kubespray/</guid>
    <pubDate>Fri, 29 Nov 2019 22:23:25 +0000</pubDate>
    <description>If you’ looking into serverless computing, you probably have bumped into Apache Openwhisk and Knative. Both are the opensource frameworks for serverless computing that allow you to deploy event-driven microservices,…</description>
    <content:encoded><![CDATA[<p class="ta-justify">If you&#8217; looking into serverless computing, you probably have bumped into Apache Openwhisk and Knative. Both are the opensource frameworks for serverless computing that allow you to deploy event-driven microservices, called functions.</p>
<p class="ta-justify">Apache Openwhisk deployment has many options, including on top of a running Kubernetes cluster. You may always use the Kubernetes deployment tools such as <a href="https://rancher.com/docs/rancher/v2.x/en/installation/ha/kubernetes-rke/" target="_blank" rel="noopener noreferrer">RKE</a> or <a href="https://github.com/kubernetes-sigs/kubespray" target="_blank" rel="noopener noreferrer">Kubespray</a> to deploy Kubernetes and later use helm charts to deploy Apache Openwhisk. I found this the most consistent way of creating Apache Openwhisk deployments for evaluation and performance analysis.</p>
<p class="ta-justify">If you used Kubespray to deploy Kubernetes,  note that starting version <a href="https://github.com/kubernetes-sigs/kubespray/releases/tag/v2.9.0" target="_blank" rel="noopener noreferrer">2.9.0</a> , Kubespray no longer supports KubeDNS, so your Kubernetes deployment will be using CoreDNS instead. This will impact your Apache Openwhisk deployment, which by default uses KubeDNS for dns resolution. When you deploy the <a href="https://github.com/apache/openwhisk-deploy-kube" target="_blank" rel="noopener noreferrer">helm chart</a> for openwhisk, you will get an error like this in the nginx pod</p>
<div class="codeblock"><pre><code>nginx: [emerg] 1#1: host not found in resolver "kube-dns.kube-system" in /etc/nginx/nginx.conf:41
nginx: [emerg] host not found in resolver "kube-dns.kube-system" in /etc/nginx/nginx.conf:41</code></pre></div>
<p>What you need to do is to change the config to use the CoreDNS by updating the k8s section in the values.yaml to use coredns, like this:</p>
<div class="codeblock"><pre><code>k8s:
  domain: cluster.local
<strong>  dns: coredns.kube-system</strong>
  persistence:
    enabled: true
    hasDefaultStorageClass: true
    explicitStorageClass: nil</code></pre></div>
<p>You will need to redeploy the helm chart.</p>
<p>Another option if you don&#8217;t want to redeploy is to edit the config map of nginx</p>
<div class="codeblock"><pre><code>kubectl edit configmap -n NAMESPACE nginx</code></pre></div>
<p>you will have to update the resolver to be like this</p>
<div class="codeblock"><pre><code>resolver coredns.kube-system;</code></pre></div>
<p>Then, you will need to restart the nginx, by scaling it down to 0 pods and then scaling it back up</p>
<div class="codeblock"><pre><code>kubectl scale deployment nginx -n NAMESPACE --replicas=0
kubectl scale deployment nginx -n NAMESPACE --replicas=1</code></pre></div>
<p>This should fix the dns resolution issue</p>
<p>Good luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>How NICs work ? a quick dive !</title>
    <link>https://mohamede.com/posts/how-nics-work-a-deep-dive/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/how-nics-work-a-deep-dive/</guid>
    <pubDate>Tue, 01 Oct 2019 13:33:18 +0000</pubDate>
    <description>I’ve written this post as a draft sometime ago, but forgot to post it. The reason I looked into it was to find out how DPDK physically works as the OS/Device level and how it bypasses the network stack. So, when you…</description>
    <content:encoded><![CDATA[<p>I&#8217;ve written this post as a draft sometime ago, but forgot to post it. The reason I looked into it was to find out how DPDK physically works as the OS/Device level and how it bypasses the network stack.</p>
<p>So, when you attach a PCIe NIC to your Linux server, you expect traffic will flow once your application send traffic. But how does it actually flow, and what components are invoked in Linux to get this traffic to flow, if you are thinking of these questions then this post is for you.</p>
<p>Let&#8217;s first discuss the main components that allow your NIC card to be recognized, registered and handled by Linux.</p>
<ul>
<li>NIC FIFO buffers: From the name, this is a FIFO Hardware buffer that your NIC has in it, the purpose is simply to put the received data somewhere before passing it to the OS. The size is determined by the vendor. One tricky part here is that it&#8217;s frequently named in the driver code as the &#8220;ring&#8221; buffer, which is true. FIFO buffers are a ring implementation, i.e. you start overriding packets if you run out of space in the buffer</li>
<li>NIC driver: The driver is by default, a kernel module. This means it lives in the kernel memory space. The driver has three core functions
<ul>
<li>Allocate RX and TX queues: Those are queues in the memory of the host server (i.e. hardware, but on your server, not your NIC). The purpose of those queues is to contain pointers to your to-be-sent/received packets. The RX and TX queues are usually refereed to as , descriptor rings. The name descriptor comes from their core purpose: containing descriptors for packets and not the packets contents.</li>
<li>Initialize the NIC: Basically reach out to the hardware NIC registers, set their values appropriately, and during operations pass them some memory addresses (contained in RX/TX descriptor rings) on where data to be sent/received will exist</li>
<li>Handling interrupts: When you driver starts, it has to register a way to communicate with the host OS, to tell it it has received packets. Also the OS has to be able to &#8220;kick&#8221; the NIC to send ready-to-be-sent packets. This is where the driver has to register &#8220;Interrupt service handlers/routines&#8221;, aka ISR. Depending on the generation/architecture of the NIC, there are multiple kinds of Interrupts supported by Linux that the NIC may implement</li>
</ul>
</li>
<li>NIC DMA Engine: Responsible for copying data in/out from the NIC FIFO buffers to your RAM. This is a physical part of the NIC</li>
<li>NAPI: New API, basically a polling mechanism that works on scheduled threads that handle new arriving data. The most common task is to get this data flowing through the network stack. NAPI relies on the concept of poll lists, that drivers register their interrupts to and then are harvested periodically, instead of continuous interrupt servicing which is CPU heavy.</li>
<li>Network Stack: Where all the OSI model happens. Except of course if you&#8217;r using DPDK, you dont want it to pass through the network stack all together</li>
</ul>
<p>A summary diagram for that is</p>
<p><picture><source type="image/webp" srcset="https://mohamede.com/assets/uploads/2019/10/nic-480.webp 480w, assets/uploads/2019/10/nic-760.webp 760w, assets/uploads/2019/10/nic-1520.webp 1520w" sizes="(max-width: 792px) 100vw, 760px"><img src="https://mohamede.com/assets/uploads/2019/10/nic.png" alt="NIC.png" width="1600" height="945" loading="lazy" decoding="async"></picture></p>
<p>Let&#8217;s follow what happens when a NIC receives a packet:</p>
<ul>
<li>First thing, the NIC will put the packet in the FIFO memory</li>
<li>Second, the NIC will use the dma engine to try to receive an RX descriptor. The RX descriptor will point to the location in memory to store the received packet. Descriptors only point to the location of the received data and do not contain the data itself</li>
<li>Thrid, once the NIC knows where to put the data. It will use the DMA engine again to write the received packet to the memory region specified in the descriptor.</li>
<li>Once the data is in the received memory region, the NIC will raise an RX interrupt to the host OS.</li>
<li>Depending on the type of enabled interrupts, the OS either stops what&#8217;s doing to handle this interrupt (CPU heaviy), or relies on a polling mechanism that works regularly to check for the new interrupts (NAPI.. aka New API) which is the default method in newer kernels</li>
<li>NAPI hand over the data in the memory region to the network stack and that&#8217;s when it goes through multiple layers of the OSI model to eventually reach a socket where your application is waiting in the user-space.</li>
</ul>
<p>A very good go-to manual is located at:</p>
<p>https://blog.packagecloud.io/eng/2016/06/22/monitoring-tuning-linux-networking-stack-receiving-data</p>
<p>Good luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>Open Infra Summit Shanghai 2019</title>
    <link>https://mohamede.com/posts/openstack-summit-shanghai-2019/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/openstack-summit-shanghai-2019/</guid>
    <pubDate>Mon, 27 May 2019 07:39:23 +0000</pubDate>
    <description>I’m pleased to serve on the programming committee of the “AI, Machine Learning and HPC” track in the upcoming Open Infra summit in Shanghai. If you plan to submit a talk and have questions, looking for advice or just…</description>
    <content:encoded><![CDATA[<p>I&#8217;m pleased to serve on the programming committee of the <strong>&#8220;AI, Machine Learning and HPC&#8221;</strong> track in the upcoming <a href="https://www.openstack.org/summit/shanghai-2019/" target="_blank" rel="noopener noreferrer">Open Infra summit in Shanghai</a>. If you plan to submit a talk and have questions, looking for advice or just want to have a chat on your proposal. I will be happy to help you with crafting the proposal.</p>
<p>My office hours are:</p>
<p>20 &#8211; 21 UTC, Mondays on <a href="https://webchat.freenode.net/" target="_blank" rel="noopener noreferrer">Freenode</a> IRC, #open-infra-summit-cfp</p>
<p>Good Luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>Our talk at the Open Infrastructure summit in Denver</title>
    <link>https://mohamede.com/posts/our-talk-at-the-open-infrastructure-summit-in-denver/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/our-talk-at-the-open-infrastructure-summit-in-denver/</guid>
    <pubDate>Wed, 08 May 2019 21:02:54 +0000</pubDate>
    <description>https://www.openstack.org/summit/denver-2019/summit-schedule/events/23030/openstack-in-the-canadian-arc-a-real-life-use-case Direct Youtube link Read More</description>
    <content:encoded><![CDATA[<p><a href="https://www.openstack.org/summit/denver-2019/summit-schedule/events/23030/openstack-in-the-canadian-arc-a-real-life-use-case" target="_blank" rel="noopener noreferrer">https://www.openstack.org/summit/denver-2019/summit-schedule/events/23030/openstack-in-the-canadian-arc-a-real-life-use-case</a></p>

<p><a href="https://www.youtube.com/watch?v=JTRoNxw92Lc" target="_blank" rel="noopener noreferrer">Direct Youtube link</a></p>


<figure class="wp-block-embed"><div class="wp-block-embed__wrapper">
<div class="embed-container"><div class="video"><iframe src="https://www.youtube-nocookie.com/embed/JTRoNxw92Lc?feature=oembed" title="Video" loading="lazy" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></div></div>
</div></figure>]]></content:encoded>
  </item>
  <item>
    <title>PCI passthrough: Type-PF, Type-VF and Type-PCI</title>
    <link>https://mohamede.com/posts/pci-passthrough-type-pf-type-vf-and-type-pci/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/pci-passthrough-type-pf-type-vf-and-type-pci/</guid>
    <pubDate>Thu, 07 Feb 2019 00:36:32 +0000</pubDate>
    <description>Passthrough has became more and more popular with time. It started initially for simple PCI device assignment to VMs and then grew to be part of high performance network realm in the Cloud such as SR-IOV, Host-level…</description>
    <content:encoded><![CDATA[<p>Passthrough has became more and more popular with time. It started initially for simple PCI device assignment to VMs and then grew to be part of high performance network realm in the Cloud such as SR-IOV, Host-level DPDK and VM-Level DPDK for NFV.</p>
<p>In Openstack, if you need to passthrough a device on your compute hosts to the VMs, you will need to specify that in the nova.conf via the passthrough_whitelist and the alias directives under the [pci] category. A typical configuration of nova.conf on the controller node will look like that</p>
<div class="codeblock"><pre><code>[pci]
alias = { "vendor_id":"1111", "product_id":"1111", "device_type":"type-PCI", "name":"a1"}
alias = { "vendor_id":"2222", "product_id":"2222", "device_type":"type-PCI", "name":"a2"}</code></pre></div>
<p>while on the compute host it, nova.conf will look like that</p>
<div class="codeblock"><pre><code>[pci]
alias = { "vendor_id":"1111", "product_id":"1111", "device_type":"type-PCI", "name":"a1"}
alias = { "vendor_id":"2222", "product_id":"2222", "device_type":"type-PCI", "name":"a2"}
passthrough_whitelist = [{"vendor_id":"1111", "product_id":"1111"}, {"vendor_id":"2222", "product_id":"2222"}]</code></pre></div>
<p>Each alias represents a device that nova-scheduler will be capable of scheduling againist using the <strong>PciPassthroughFilter</strong> filter. The more devices you want to pass through, the more alias lines you will have to create.</p>
<p>Alias syntax is quite self explanatory. <strong>vendor_id</strong> is unique for the device vendor, <strong>product_id</strong> is unique per device, <strong>name</strong> is an identifier that you specify of this device. Both vendor_id and product_id can be obtained via the command</p>
<div class="codeblock"><pre><code>lspci -nnn</code></pre></div>
<p>You can deduce the vendor and product ids from the output as follows</p>
<div class="codeblock"><pre><code>000a:00:00.0 PCI bridge [0000]: Host Bridge  [1111:2222]</code></pre></div>
<p>In this case, the vendor_id is 1111 and the product_id is 2222</p>
<p>But how about <strong>device_type </strong>in the alias definition ? . Well device_type can be one of three values: type-PCI, type-PF and type-VF</p>
<p><strong>type-PCI</strong> is the most generic. What it does is pass-through the PCI card to the guest VM through the following mechanism:</p>
<ul>
<li>IOMMU/VT-d will be used for memory mapping and isolation, such that the Guest OS can access the memory structures of the PCI device</li>
<li>No vendor driver will be loaded for the PCI device in the compute host OS</li>
<li>The Guest VM will handle the device directly using the vendor driver</li>
</ul>
<p>When a PCI device gets attached to a qemu-kvm instance, the libvirt definition for that instance will include a hostdev for that device, for example:</p>
<div>
<div class="codeblock"><pre><code>   &lt;hostdev mode='subsystem' type='pci' managed='yes'&gt;
     &lt;source&gt;
        &lt;address domain='0x1111' bus='0x11' slot='0x11' function='0x1'/&gt;
      &lt;/source&gt;      
        &lt;address type='pci' domain='0x1111' bus='0x11' slot='0x1' function='0x0'/&gt;
    &lt;/hostdev&gt;</code></pre></div>
</div>
<p>The next two types are more interesting. They originated for SR-IOV capable devices, where the notion of Physical function &#8220;PF&#8221; and Virtual Functions &#8220;VF&#8221;. There&#8217;s a core difference with those two types than the type-PCI which is</p>
<ul>
<li>A PF driver is loaded for the SR-IOV device in the compute-host OS.</li>
</ul>
<p>Let&#8217;s explain what the difference between type-VF and type-PF is, we will start with VFs first:</p>
<p><strong>type-VF</strong> allows you to pass a Virtual Function, which is a lightweight PCIe device that has its own RX/TX queues in case of network devices. Your VM will be able to use the VF driver, provided by the vendor, to access the VF and deal with it as a regular device for IO. VFs generally have the same vendor_id as the hardware device vendor_id, but with a different product_id specified for the VFs.</p>
<p><strong>type-PF</strong> on the other hand refers to a fully capable PCIe device, that can control the physical functions of an SR-IOV capable device, including the configuration of the Virtual functions. type-PF allows you to passthrough the PF to be controlled by the VMs. This is sometimes useful in NFV use-cases.</p>
<p>A simplified layout of PF/VF looks like this</p>
<p class="ta-left"><picture><source type="image/webp" srcset="https://mohamede.com/assets/uploads/2019/02/sriov-kernel-2-480.webp 480w, assets/uploads/2019/02/sriov-kernel-2-760.webp 760w, assets/uploads/2019/02/sriov-kernel-2-1150.webp 1150w" sizes="(max-width: 792px) 100vw, 760px"><img src="https://mohamede.com/assets/uploads/2019/02/sriov-kernel-2.png" alt="SRIOV-KERNEL (2).png" width="1150" height="1600" loading="lazy" decoding="async"></picture></p>
<p>PF driver is used to configure the SR-IOV functionality and partition the device into virtual functions accessed by the VM in userspace</p>
<p>A nice feature about nova-compute is that it does print out the Final resource view, which contains specifics of the passthroughed devices. It will look like that in the case of a PF passthrough</p>
<div class="codeblock"><pre><code>Final resource view: pci_stats=[PciDevicePool(count=2,numa_node=0,product_id='2222',tags={dev_type='type-PF'},vendor_id='1111')]</code></pre></div>
<p>Which says there&#8217;r two devices in numa cell 0 with the specified vendor_id and product_id that are available for passthrough</p>
<p>In the case of VF passthrough:</p>
<div class="codeblock"><pre><code>Final resource view: pci_stats=[PciDevicePool(count=1,numa_node=0,product_id='3333',tags={dev_type='type-VF'},vendor_id='1111')]</code></pre></div>
<p>In this case there&#8217;s only one VF with vendor_id 1111 and product_id 3333 that&#8217;s ready to be passthroughed on numa cell 0</p>
<p>The blueprint of PF passthrough type is here if you&#8217;r interested</p>
<p>https://blueprints.launchpad.net/nova/+spec/sriov-physical-function-passthrough</p>
<p>Good Luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>VNI Ranges: What do they do ?</title>
    <link>https://mohamede.com/posts/vni-ranges-what-do-they-do/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/vni-ranges-what-do-they-do/</guid>
    <pubDate>Fri, 23 Nov 2018 04:59:00 +0000</pubDate>
    <description>Deployment tools for Openstack have become very popular, including the very well known Openstack-Ansible. It makes deploying a Cloud an easy task, at the expense of losing access to the insights of “Behind the Scenes”…</description>
    <content:encoded><![CDATA[<p class="wp-block-paragraph">Deployment tools for Openstack have become very popular, including the very well known Openstack-Ansible. It makes deploying a Cloud an easy task, at the expense of losing access to the insights of &#8220;Behind the Scenes&#8221; of your your Cloud deployment. If you have had to configure neutron manually, you would have come across the following section in the ml2 configuration</p>

<div class="codeblock"><pre><code>[ml2_type_vxlan] # <br />(ListOpt) Comma-separated list of &lt;vni_min&gt;:&lt;vni_max&gt; <br />tuples enumerating # ranges of VXLAN VNI IDs that are available for <br />tenant network allocation. # <br /># vni_ranges =</code></pre></div>

<p class="wp-block-paragraph">You probably have set it to a range, similar to 10:100 or 10:300 and so on</p>

<p class="wp-block-paragraph">But what does this configuration mean ?</p>

<p class="wp-block-paragraph">When you configure neutron to use VXLAN as the segmentation network, each tenant network gets assigned a Virtual Network Identifier &#8220;VNI&#8221;. VNIs are numeric values that you specify their range with the vni_ranges parameter. </p>

<p class="wp-block-paragraph">An advantage of having control on this parameter is that you can specify the maximum number of VXLANs that the ml2 agent can use. Although this seems like an advantage, it can also be a disadvantage in a dynamic environment as you can run into situations where your networks can not be created because all allowed VNIs are consumed. If that&#8217;s the case, you will get an error similar to the following in neutron logs</p>

<div class="codeblock"><pre><code>Unable to create the network. No tenant network is available for allocation."</code></pre></div>

<p class="wp-block-paragraph">If you get that error, it means you need to increase the available ranges and restart the services to get it updated. </p><p>Best of Luck ! </p>]]></content:encoded>
  </item>
  <item>
    <title>Port security in Openstack</title>
    <link>https://mohamede.com/posts/port-security-in-openstack/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/port-security-in-openstack/</guid>
    <pubDate>Fri, 28 Sep 2018 21:16:51 +0000</pubDate>
    <description>Openstack Neutron provides by default some protections for your VMs’ communications, those protections verify that VMs can not impersonate other VMs. You can easily see how it does that by checking the flow rules in an…</description>
    <content:encoded><![CDATA[<p>Openstack Neutron provides by default some protections for your VMs&#8217; communications, those protections verify that VMs can not impersonate other VMs. You can easily see how it does that by checking the flow rules in an OVS deployment using:</p>
<div class="codeblock"><pre><code>ovs-ofctl dump-flows br-int</code></pre></div>
<p>If you look for a certain qvo port (or the port number, depending on the deployment), this will show the following lines</p>
<div class="codeblock"><pre><code>table=24, n_packets=1234, n_bytes=1234, priority=2,arp,in_port="qvo",arp_spa=10.10.10.10 actions=resubmit(,25)
table=24, n_packets=1234, n_bytes=1234, priority=0 actions=drop</code></pre></div>
<p>Table 24 by default will drop all the packets originated from a VM unless they are resubmitted to table 25. The criteria for submitting to table 25 is simple: That the source IP for this traffic is the one that has been assigned to that VM, if not it will drop the packet at the end of table 24</p>
<p>In addition , there&#8217;s a protection from changing the MAC address of the interface, it&#8217;s implemented via the following rule</p>
<div class="codeblock"><pre><code>table=25, n_packets=1234, n_bytes=1234, priority=2,in_port="qvo",dl_src=aa:aa:aa:aa:aa:aa actions=resubmit(,60)</code></pre></div>
<p>which basically compares the source MAC address of the packet with the expected MAC address of the VM.</p>
<p>In some use cases, you may want to drop this protection, it can be done using</p>
<div class="codeblock"><pre><code>neutron port-update $PORT_ID --port-security-enabled=false</code></pre></div>
<p>This will ensure there&#8217;s no openflow rules in br-int that will drop your packets if they don&#8217;t adhere to the MAC/IP requirements</p>
<p>Good Luck !</p>]]></content:encoded>
  </item>
  <item>
    <title>Migrating VMs with attached RBDs</title>
    <link>https://mohamede.com/posts/migrating-vms-with-attached-rbds/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/migrating-vms-with-attached-rbds/</guid>
    <pubDate>Fri, 29 Jun 2018 23:14:08 +0000</pubDate>
    <description>From the title, this is obviously a very common scenario that you may want to do. One thing that we rarely think about though is “backends” for the attached volumes when we create volumes. When you create a volume, the…</description>
    <content:encoded><![CDATA[<p>From the title, this is obviously a very common scenario that you may want to do. One thing that we rarely think about though is &#8220;backends&#8221; for the attached volumes when we create volumes.</p>
<p>When you create a volume, the volume is created on a cinder backend and kept attached to this backend until it&#8217;s deleted , or migrated to another backend. The backends are defined in cinder configuration and are provided by your host(s) running the cinder-volume service. To find your backends, run the following command</p>
<div class="codeblock"><pre><code>cinder get-pools</code></pre></div>
<p>When you attach the volume to a VM,  the volume keeps its backend. It relies on this backend to do any operations to that volume. This includes migrating the VM from a host to another.</p>
<p>You may run into a scenario where you get this error when trying to migrate a VM with attached RBD</p>
<div class="codeblock"><pre><code> CinderConnectionFailed: Connection to cinder host failed: Unable to establish connection to</code></pre></div>
<p>But when you go and check, Cinder is working correctly. You are able to create new volumes and attach them to instances. But a particular VM is unable to migrate. You may find also you&#8217;r unable to snapshot the volume attached to the VM. The thing to check for here is the RBD backend of the volume</p>
<p>You can find this using</p>
<div class="codeblock"><pre><code>cinder show VOLUME_ID</code></pre></div>
<p>this will show you alot of details on the volume including the following attribute</p>
<div class="codeblock"><pre><code>| os-vol-host-attr:host | HOSTNAME@ceph#RBD |</code></pre></div>
<p>HOSTNAME will likely be &#8220;one&#8221; of your controllers. You will need to go and check that cinder-volume service is running correctly on that controller. If it&#8217;s down, you can&#8217;t operate that volume for anything (snapshots, attach/detach and migrate)</p>
<p>If you&#8217;ve lost your controller forever, or you were testing a new backend that no longer exists, then you might want to migrate the volume from the dead backend. This is detailed in the following manual</p>
<div class="codeblock"><pre><code>https://docs.openstack.org/cinder/pike/admin/blockstorage-volume-migration.html</code></pre></div>
<p>Happy VM migrations !</p>]]></content:encoded>
  </item>
  <item>
    <title>My talk at CANHEIT-TECC 2018</title>
    <link>https://mohamede.com/posts/my-talk-at-canheit-2018/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/my-talk-at-canheit-2018/</guid>
    <pubDate>Thu, 21 Jun 2018 00:50:21 +0000</pubDate>
    <description>http://programme.exordo.com/canheit-tecc2018/delegates/presentation/30/ Read More</description>
    <content:encoded><![CDATA[<p><a href="http://programme.exordo.com/canheit-tecc2018/delegates/presentation/30/" target="_blank" rel="noopener noreferrer">http://programme.exordo.com/canheit-tecc2018/delegates/presentation/30/</a></p>]]></content:encoded>
  </item>
  <item>
    <title>My talk at the Openstack Summit in Vancouver</title>
    <link>https://mohamede.com/posts/my-talk-at-the-openstack-summit-in-vancouver/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/my-talk-at-the-openstack-summit-in-vancouver/</guid>
    <pubDate>Tue, 22 May 2018 19:37:28 +0000</pubDate>
    <description>https://www.openstack.org/summit/vancouver-2018/summit-schedule/events/20630/dpdk-putting-neutron-in-the-fast-lane Direct Youtube Video Read More</description>
    <content:encoded><![CDATA[<p><a href="https://www.openstack.org/summit/vancouver-2018/summit-schedule/events/20630/dpdk-putting-neutron-in-the-fast-lane" target="_blank" rel="noopener noreferrer">https://www.openstack.org/summit/vancouver-2018/summit-schedule/events/20630/dpdk-putting-neutron-in-the-fast-lane</a></p>
<p><a href="https://www.youtube.com/watch?v=vJSNfCzj15w" target="_blank" rel="noopener noreferrer">Direct Youtube Video</a></p>


<figure class="wp-block-embed"><div class="wp-block-embed__wrapper">
<div class="embed-container"><div class="video"><iframe src="https://www.youtube-nocookie.com/embed/vJSNfCzj15w?feature=oembed" title="Video" loading="lazy" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen></iframe></div></div>
</div></figure>]]></content:encoded>
  </item>
  <item>
    <title>Quota usage refresh in Openstack</title>
    <link>https://mohamede.com/posts/quota-usage-refresh-in-openstack/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/quota-usage-refresh-in-openstack/</guid>
    <pubDate>Fri, 16 Feb 2018 03:44:10 +0000</pubDate>
    <description>Openstack stores quota usage for tenants in the database in quota_usages table. Nova and cinder have by default their own separate databases and in each database you get a new quota_usages table. The structure of the…</description>
    <content:encoded><![CDATA[<p>Openstack stores quota usage for tenants in the database in quota_usages table. Nova and cinder have by default their own separate databases and in each database you get a new quota_usages table.</p>
<p>The structure of the quota_usages table is as follows</p>
<div class="codeblock"><pre><code>+---------------+--------------+------+-----+---------+----------------+
| Field | Type | Null | Key | Default | Extra |
+---------------+--------------+------+-----+---------+----------------+
| created_at | datetime | YES | | NULL | |
| updated_at | datetime | YES | | NULL | |
| deleted_at | datetime | YES | | NULL | |
| id | int(11) | NO | PRI | NULL | auto_increment |
| project_id | varchar(255) | YES | MUL | NULL | |
| resource | varchar(255) | NO | | NULL | |
| in_use | int(11) | NO | | NULL | |
| reserved | int(11) | NO | | NULL | |
| until_refresh | int(11) | YES | | NULL | |
| deleted | int(11) | YES | | NULL | |
| user_id | varchar(255) | YES | MUL | NULL | |
+---------------+--------------+------+-----+---------+----------------+</code></pre></div>

<p>Remember that quotas are managed per project, so in this table project_id is your navigating key. Project IDs are used to identify projects. For a particular project, you can retrieve the project ID using</p>
<div class="codeblock"><pre><code>openstack project list | grep $PROJECT_NAME</code></pre></div>
<p>The other interesting fields in the quota_usages table are</p>
<p>resource: for example in nova, it can be &#8220;instance, ram, cores, security_groups&#8221;</p>
<p>in_use : This is the amount per resource that Openstack &#8220;thinks&#8221; that project is using</p>
<p>Occasionally the in_use field is not updated properly and you might find yourself in a situation where openstack is reporting usage that doesn&#8217;t exist. You have two options at this point</p>
<ul>
<li>Use the nova-manage project quota_usage_refresh command to try to refresh the quota for a specific project. The syntax is something like</li>
</ul>
<div class="codeblock"><pre><code>nova-manage project quota_usage_refresh --project PROJECT_ID --user USER_ID --key cores</code></pre></div>
<ul>
<li>If that doesn&#8217;t help, you may have to update the MySQL database using the update statement. You will need to restart the respective service after that to see the change</li>
</ul>]]></content:encoded>
  </item>
  <item>
    <title>Glance and CEPH backend</title>
    <link>https://mohamede.com/posts/glance-and-ceph-backend/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/glance-and-ceph-backend/</guid>
    <pubDate>Fri, 05 Jan 2018 03:44:53 +0000</pubDate>
    <description>Using CEPH as a backend for glance images has slowly become the default deployment methodology in many production deployments. It is usually as easy as creating a new pool in ceph ( glance pool) and creating a user to…</description>
    <content:encoded><![CDATA[<p>Using CEPH as a backend for glance images has slowly become the default deployment methodology in many production deployments. It is usually as easy as creating a new pool in ceph ( glance pool) and creating a user to be associated with glance. The glance CEPH user will normally authenticate using cephx and store images and snapshots in CEPH.</p>
<p>The configuration in glance-api.conf looks something like this on the controller/s</p>
<div class="codeblock"><pre><code>[glance_store]
stores = glance.store.rbd.Store
default_store = rbd
rbd_store_pool = POOLNAME
rbd_store_user = USERNAME
rbd_store_ceph_conf = /etc/ceph/ceph.conf</code></pre></div>
<p>then in the default location of the CEPH keyrings &#8220;/etc/ceph&#8221;, you will need to add the keyring for the CEPH user associated with glance.</p>
<p>On CEPH nodes, don&#8217;t forget to grant permissions to the glance user to the glance pool. The permissions need to be read/write such that it can create new images and read the existing ones. The default command to create a new user and grant read/write permissions to the pool is:</p>
<div class="codeblock"><pre><code>ceph auth get-or-create client.user mon ‘allow r’ osd ‘allow class-read object_prefix rbd_children, allow rwx pool=glancepool’ -o /etc/ceph/ceph.client.images.keyring</code></pre></div>
<p>If after you create the pool and configure glance-api and the keyring properly on the controller node you get something like this in nova-conductor.log when provisioning new VMs</p>
<div class="codeblock"><pre><code>WARNING nova.scheduler.utils [req-a3c6f93e-484a-43e0-9e73-5bbdc451b2c6   - - -] Failed to compute_task_build_instances: Exceeded maximum number of retries. Exceeded max scheduling attempts 3 for instance ID. Last exception: HTTPInternalServerError (HTTP 500)
ERROR nova.scheduler.utils [req-a3c6f93e-484a-43e0-9e73-5bbdc451b2c6  - - -] [instance: ID] Error from last host: HOST (node HOST): [u'Traceback (most recent call last):\n', u' File "/usr/lib/python2.7/site-packages/nova/compute/manager.py", line 1780, in _do_build_and_run_instance\n filter_properties)\n', u' File "/usr/lib/python2.7/site-packages/nova/compute/manager.py", line 2016, in _build_and_run_instance\n instance_uuid=instance.uuid, reason=six.text_type(e))\n', u'RescheduledException: Build of instance ID was re-scheduled: HTTPInternalServerError (HTTP 500)\n']</code></pre></div>
<p>check the glance logs as well, you will most likely find a 500 error in glance logs (api.log)</p>
<div class="codeblock"><pre><code>INFO eventlet.wsgi.server [req-a3c6f93e-484a-43e0-9e73-5bbdc451b2c6 ] IP "GET /v2/images/2IDfile HTTP/1.1" <span class="hl-red"><strong>500</strong></span> 139 0.182237</code></pre></div>
<p>This error is not very indicative, however it means that you would want to check the permissions on the keyring file for the glance user in the /etc/ceph and make sure that the user running glance (by default &#8220;glance&#8221; ) has read permissions to it at least</p>
<p>For example</p>
<div class="codeblock"><pre><code>ls -l /etc/ceph/ceph.client.glancepool.keyring 
-r--r----- 1 glance glance 64 Oct 27 11:12 /etc/ceph/ceph.client.glancepool.keyring</code></pre></div>
<p>This permission is the minimum that glance can access the glance pool in CEPH</p>
<p>Have fun !</p>]]></content:encoded>
  </item>
  <item>
    <title>VM Cold migrations/resizing in openstack</title>
    <link>https://mohamede.com/posts/vm-cold-migrations-in-openstack/</link>
    <guid isPermaLink="true">https://mohamede.com/posts/vm-cold-migrations-in-openstack/</guid>
    <pubDate>Thu, 19 Oct 2017 23:36:50 +0000</pubDate>
    <description>Cold migrations are an integral piece of any QEMU/KVM deployment. It’s cold or “non-live” as you have to power down the VM, move it to the new host and power it back up. Openstack follows the same procedure when it…</description>
    <content:encoded><![CDATA[<p>Cold migrations are an integral piece of any QEMU/KVM deployment. It&#8217;s cold or &#8220;non-live&#8221; as you have to power down the VM, move it to the new host and power it back up. Openstack follows the same procedure when it comes to migrating VMs.</p>
<p>Cold migrations in Openstack are done via the user running the openstack-nova-compute process. This user is &#8220;nova&#8221; in most cases. In order to allow cold migrations, the user &#8220;nova&#8221; has to be able to ssh, password-less, to the other compute hosts in the environment. This is why this part of openstack configuration is needed</p>
<p>https://docs.openstack.org/nova/pike/admin/ssh-configuration.html#cli-os-migrate-cfg-ssh</p>
<p>The basic steps for cold migrations are as follows</p>
<ul>
<li>An admin user initiates the migration using either &#8220;openstack server migrate&#8221; , &#8220;nova migrate&#8221; or the dashboard</li>
<li>nova-scheduler tries to identify a target host based on the scheduler configuration the VM specs</li>
<li>openstack-nova-compute uses the &#8220;nova&#8221; user to ssh to the target host, create the needed /var/lib/nova/instances directories, copy the VM definitions and the disk over to the target system</li>
<li>nova-scheduler now starts the new VM on the new host</li>
</ul>
<p>A common error that you might see if you have the ssh keys misconfigured is a variation of the following error.</p>
<div class="codeblock"><pre><code>2017-10-17 16:38:27.149 8053 ERROR oslo_messaging.rpc.server ResizeError: Resize error: not able to execute ssh command: Unexpected error while running command.
2017-10-17 16:38:27.149 8053 ERROR oslo_messaging.rpc.server Command: ssh -o BatchMode=yes IP mkdir -p /var/lib/nova/instances/ID
2017-10-17 16:38:27.149 8053 ERROR oslo_messaging.rpc.server Exit code: 255
2017-10-17 16:38:27.149 8053 ERROR oslo_messaging.rpc.server Stdout: u''
2017-10-17 16:38:27.149 8053 ERROR oslo_messaging.rpc.server Stderr: u'Host key verification failed.\r\n'</code></pre></div>
<p>To test it, try switching to the nova user on the source compute host using &#8220;su &#8211; nova&#8221; and ssh to the target host.</p>
<div class="codeblock"><pre><code>su - nova
ssh target-host</code></pre></div>
<p>If you get a similar message</p>
<div class="codeblock"><pre><code>Warning: Permanently added the RSA host key for IP address 'IP' to the list of known hosts.</code></pre></div>
<p>This indicates that: although you have setup the password-less ssh properly, the host key of the target compute host was not &#8220;yet&#8221; trusted on the source compute host.  This will mostly happen if you copied the keys manually and did not use the ssh-copy-id command</p>]]></content:encoded>
  </item>
</channel>
</rss>
