9/23/2020

Understanding the output of "show qos interface eth#"

In this post, I like to explain the output of the EOS command - "show qos interface eth#" based on my understanding from EOS document. 

wa461.00:58:25#sh qos interfaces e17/1
Ethernet17/1:
   Trust Mode: DSCP
   Default COS: 0
   Default DSCP: 0

   Port shaping rate: disabled
   Burst-size: disabled

  Tx    Bandwidth         Shape Rate               Burst-Size          Priority   ECN/WRED
 Queue  (percent)          (units)                  (units)
 ------------------------------------------------------------------------------------------
   7      - / -       - / -          ( - )           -  /  -           SP / SP       D
   6      - / -       - / -          ( - )           -  /  -           SP / SP       D
   5      - / -       - / -          ( - )           -  /  -           SP / SP       D
   4      - / -       - / -          ( - )           -  /  -           SP / SP       D
   3      - / -       - / -          ( - )           -  /  -           SP / SP       D
   2     20 / 20    1.2 / 1.0        (Gbps)     2048 KB / 2048 KB      RR / RR       D
   1     30 / 30      - / -          ( - )           -  /  -           RR / SP       D
   0     50 / 50      - / -          ( - )           -  /  -           RR / SP       D

Note: Values are displayed as Operational/Configured
Legend:
RR -> Round Robin
SP -> Strict Priority
 - -> Not Applicable / Not Configured
 % -> Percentage of line rate

  • Values are displayed as Operational/Configured, like RR/SP which means this Q is configured as strict priority but operational as round-robin. 
  • If one queue is configured as no priority (RR), then all the lower queues are changed to RR
    • In this example, Q 2 is RR, then 0 and 1 are automatically changed to RR. 
    • And Q 0 and 1 are RR/SP, which means their configuration are SP by default, but operational mode is RR.
  • If both interface and tx-queue are configured with shape, which is effective? 
    • From EOS manual chapter 27.5 - Enabling port shaping on an FM6000 interface disables queue shaping internally. Disabling port shaping restores queue shaping as specified in running-config.
interface Ethernet17/1
   speed forced 10000full
   !
   tx-queue 0
      bandwidth percent 50
   !
   tx-queue 1
      bandwidth percent 30
   !
   tx-queue 2
      no priority
      bandwidth percent 20
      shape rate 1000000
  • Bandwidth vs shape.  
    • Bandwidth% is the b/w percent this RR queue can get. Says the above configuration:
      • In the sample below, the interface 17/1 is 10 Gbps interface 
      • Q3-7 are the strict priority and, say use total 2 Gbps traffic, which left 8Gbps for Q0-2
      • The tx-Q 2 can have 20% of left-over capacity which is 1.6Gbps
      • But the shape rate is 1.2Gbps
      • So the maximum throughput of tx-Q 2 is 1.2 Gbps, even it is assigned with 1.6Gbps.

    EOS: A simple Qos design example

    This article - "A Simple Quality of Service Design Example" is a very good starting point for understanding the EOS Qos architecture and starting a Qos design. 

    Some points:

    • 3 ways in the ingress points to map packets to Tx queues:
      • qos cos trust + cos-tc map
      • qos dscp trust + dscp-tc map
      • service-policy + policy-map
    • 3 big categories of traffic:
      • network-control = control plane
      • latency/jitter sensitive traffic
      • best-efforts = scavenger traffic
    • qos profile = 
      • policy-map for input
      • tx-queue set for output
    • "no priority" in a tx-queue, all lower queues become RR

    9/22/2020

    Support for CPU traffic policy

    The default EOS Copp only provides protocol level traffic policy, for example, the maximum throughput of bgp traffic destined to the cpu. But there is no granularity of source address. So this feature - Support for CPU traffic policy is for this purpose. 


    Set DSCP value for CPU outbound packets

    In Arista EOS, the following protocol packets are able to set a DSCP value other than the default value 0:

    Notes:

    • The setting must be done individually and under the protocol section
    • hostname is not supported. 
    For example, setting all locally originated CP packets to DSCP value nc1/cs6/110000/63:

    logging qos dscp 48
    !
    dns qos dscp 48
    !
    ntp qos dscp 48
    !
    traceroute qos dscp 48
    !
    sflow qos dscp 48
    snmp-server qos dscp 48
    radius-server qos dscp 48
    tacacs-server qos dscp 48
    management ssh
       qos dscp 48
    !
    router ospf general
       qos dscp 48


    9/13/2020

    EOS: % Not supported when show bgp summary

    If you see the error message with EOS command - show bgp <AF> summary, it is probably caused the routing mode. To be more specific, you are probably running ribd mode and the CLI - "show bgp <AF> summary" is only supported in multi-agent mode

    ghs259#show bgp ipv6 unicast summary
    % Not supported

    ghs259-CIN-DPA2.23:32:18#show ip route summary
    Operating routing protocol model: ribd
    Configured routing protocol model: multi-agent (will apply after next reboot)

    8/19/2020

    Arista EOS - BGP Selective Route Download

     In this post, I will share my experience with a relatively old (was released back in 2015) but very useful Arista EOS feature - BGP Selective Route Download (SRD)

    The use cases are quite straightforward:

    • Program the necessary routes on the routers with small hardware resources. In the above TOI link, only 30K prefixes of 520K (back in 2015) cover 99% traffic. The left small traffic can be directed by the default route. 
    • Another useful case (for me) is to control what routes be programmed, or even not installed at all. At meanwhile the BGP runs transparently, which processes, receives and advertises the BGP prefixes. A good example is the RR which is not in the data path.  Or hardness router in the lab, it just sends bgp updates. The traffic can be handled by a couple of static routes. 
    In the below example, BGP only installs /24 IPv4 routes within 110.0.0/8 range and /64 IPv6 routes in 2000:110:1::/48. 

    router bgp 65501
       bgp route install-map part-peer-v46
    !
    route-map part-peer-v46 permit 10
       match ip address prefix-list part-peer-route
    !
    route-map part-peer-v46 permit 20
       match ipv6 address prefix-list part-peer-route-v6
    !
    ip prefix-list part-peer-route seq 10 permit 110.0.0.0/8 eq 24
    !
    ipv6 prefix-list part-peer-route-v6
        seq 10 permit 2000:110:1::/48 eq 64

    bn309#show ip route summary
    ...
    VRF: default
       Route Source                                Number Of Routes
    ------------------------------------- -------------------------
    ...
       ospfv3                                                     0
       bgp                                                     1814
         External: 1814 Internal: 0
    ...
       Total Routes                                            1871

    Number of routes per mask-length:
       /8: 2         /12: 2        /16: 1        /24: 1816     /25: 1
       /30: 7        /32: 42

    bn309#show ip bgp installed | egrep '^ \* ' | wc -l
    1816

    At this time actually, this router receives/accepts over 1.4M prefixes. 

    bn309#show ip bgp summary
    BGP summary information for VRF default
    Router identifier 192.168.230.2, local AS number 65501
    Neighbor Status Codes: m - Under maintenance
      Description              Neighbor         V  AS           MsgRcvd   MsgSent  InQ OutQ  Up/Down State   PfxRcd PfxAcc
      IpTransit#1-7504         100.101.1.1      4  12083          44303    119708    0    0 01:08:12 Estab   690049 690049
      Local-Aris-Simu          192.168.230.1    4  65510         129394         5    0    0 01:03:57 Estab   819857 819857

    As of August 2020, this feature is only supported on RIBD (so multi-agent mode doesn't work)

    8/10/2020

    EOS: alias to sum up total num of received bgp prefixes

    EOS-R1#show ip bgp neighbors 
    BGP neighbor is 100.101.1.2, remote AS 100, external link
      Prefix Statistics:
                                       Sent      Rcvd     Best Paths     Best ECMP Paths
        IPv4 Unicast:                688000    818876         808929                   0
        IPv6 Unicast:                     0         0              0                   0

    If we like to know the total number of rcvd prefix from all bgp peers, here is the alias command could be useful

    alias totbgp show ip bgp neighbors | grep "IPv4 Unicast: \s\s" | awk  '{s+=$4}END{print s}'


    EOS-R1#totbgp
    6001256

    7/27/2020

    EOS: sum up and compare in/egress throughput

    Sometimes you want to compare the ingress/egress throughput on a particular router to see if any possible traffic loss (of course, the loss should be large enough like 3% more).  On EOS, srnz (alias srnz Show interface counters rates | nz) is a good alias. But if incoming or outgoing on multiple ports, you have to sum up and compare.

    Here is a couple of useful tips and commands.

    bn309#srnz
    Port      Name        Intvl   In Mbps      %  In Kpps  Out Mbps      % Out Kpps
    Et9/1/1   ixia:LC8     0:05       0.0   0.0%        0   13435.8  35.0%     3543
    Et9/2/1   ixia:LC8     0:05       0.0   0.0%        0   13435.3  35.0%     3543
    ...
    Et11/6/1  ixia:LC7     0:05   13433.7  35.0%     3543       0.0   0.0%        0
    Et11/11/1 ixia:LC7     0:05   13433.7  35.0%     3543       0.0   0.0%        0
    Et11/12/1 ixia:LC7     0:05   13434.6  35.0%     3543       0.0   0.0%        0
    Et11/13/1 ixia:LC7     0:05   13432.3  35.0%     3543       0.0   0.0%        0
    Et11/14/1 ixia:LC7     0:05   13434.4  35.0%     3543       0.0   0.0%        0
    Et11/15/1 ixia:LC7     0:05   13434.3  35.0%     3543       0.0   0.0%        0
    Et11/16/1 ixia:LC7     0:05   13433.9  35.0%     3543       0.0   0.0%        0

    In the above example, you want to compare ingress from ixia:LC7 and egress of ixia:LC8

    bn309#srnz | grep LC7 | awk '{s+=$4}END{print s}'
    161206  <<< ingress
    bn309#srnz | grep LC8 | awk '{s+=$7}END{print s}'
    161198  <<< egress

    6/30/2020

    "no-internal-vlan" error for routed interfaces

    Creating several L3 routed port-channels and sub-interfaces, but they failed to come up with errdisabled status. The output "show interface status err" displays the following reasons:

    yo411.16:22:08(config-if-Po1201)#show int status errdisabled
       Port           Name             Status         Reason
    -------------- ---------------- ----------------- ---------------------
       Et3/12/1                        errdisabled    port-channel-shutdown
       Et4/12/1                        errdisabled    port-channel-shutdown
       Po1201.3                        errdisabled    no-internal-vlan
       Po1201.2                        errdisabled    no-internal-vlan
       Po1201                          errdisabled    no-internal-vlan


    Basically, the "port-ch-shutdown" error was caused by the Po1201 being down. Checked the EOS document, the system will reserve an internal VLAN for any "no switchport" interfaces. And the internal VLAN ranges start from 1006 to 4094 (ref: EOS Manual section 19.4.3)

    yo411.16:22:59(config-if-Po1201)#show vlan internal usage
    1006  Port-Channel1900.4002
    1007  Ethernet3/36/3
    1008  Port-Channel1900
    1009  Ethernet3/36/1
    1010  Port-Channel1900.4003
    1011  Ethernet3/36/4
    1012  Ethernet3/36/2
    1013  Ethernet3/36/4.2
    1014  Ethernet3/36/4.3


    And internal VLAN assignment stops at 1015. 

    yo411.16:23:23(config-if-Po1201)#sh vlan 1015
    VLAN  Name                             Status    Ports
    ----- -------------------------------- --------- -------------------------------
    1015  VLAN1015                         suspended

    So the root cause is that there is an accidental configuration of vlan 1015 with a suspended state, and this blocks the internal VLAN assignment. 

    yo411.16:24:19(config)#no vlan 1006 - 1099
    yo411.16:24:36(config)#sh int status errdisabled

    After removing the VLAN configuration, there is no internal-VLAN error anymore. 

    yo411.16:24:47(config)#show vlan internal usage
    ...
    1015  Port-Channel1201.2
    1016  Port-Channel1201.3
    1017  Port-Channel1201


    Another way is to specify the internal VLAN range to an unused space (ref: EOS manual section 21.3

    yo411(config)# vlan internal order descending range 4000 4094

    5/31/2020

    Arista EVPN VXLAN Configuration Example (3c) - Single-homing, L3 EVPN, Symmetric IRB

    One of the purposes of symmetric IRB is to address the scale issue of asymmetric IRB solution. And here is the list of differences compared with asymmetric IRB:
    • VTEPs only need to hold the VLANs and SVIs of directly connected subnets
    • An intermediate IP-VRF to carry the remote subnets

    5/30/2020

    Arista EVPN VXLAN Configuration Example (3b) - Single-homing, L3 EVPN, Asymmetric IRB

    To solve the sub-optimal routing pattern in the solution of centralized routing, the IRB EVPN Draft proposes 2 solutions, asymmetric IRB and symmetric IRB. Because the local VTEP does both inter-VLAN routing and intra-VLAN switching, it is called IRB (Integrated Routing and Bridging).

    The asymmetric IRB is illustrated as below



    Explanations:
    • VTEP on has 1 directly connected VLAN:
      • VTEP1 - VLAN 641
      • VTEP2 - VLAN 642
    • But the VTEPs must have
      • 2 x SVI interface, VLAN 641 and 642
      • 2 x VLANs under MAC VRF
      • 2 x VLAN/VNI bindings under Vxlan interfce
    • Routing is performed on the ingress VTEP, and egress VTEP only decapsulates the Vxlan header and switches into destination VLANs. 
    • The returning traffic does the same, so routing is done on different VTEPs, hence the term Asymmetric IRB
    • Advantage:
      • Optimal traffic path and no traffic trombone 
    • Disadvantage:
      • VTEPs must have all SVIs and VLANs configured, even not locally connected. 
      • That means ALL VTEPs hold ALL MAC and ARP of hosts for source and destination VLANs. 
      • So the scale is the biggest issue. To make things worse, TOR devices normally don't much high capacity.  
    From the below output, the VTEP1 has 6 ARP entries, 3 local VLANs and 3 remote VLANs

    snp261-eVtep1.23:11:00#sh arp
    Address         Age (sec)  Hardware Addr   Interface
    160.64.1.101      0:02:54  444c.a8a5.1140  Vlan641, Ethernet78
    160.64.1.102      0:01:08  444c.a8a5.1140  Vlan641, Ethernet78
    160.64.1.103      0:01:04  444c.a8a5.1140  Vlan641, Ethernet78
    160.64.2.201            -  444c.a8a5.1141  Vlan642, Vxlan1
    160.64.2.202            -  444c.a8a5.1141  Vlan642, Vxlan1
    160.64.2.203            -  444c.a8a5.1141  Vlan642, Vxlan1

    Arista EVPN VXLAN Configuration Example (3a) - Single-homing, L3 EVPN, Centralized Routing

    In traditional DC design, the most common inter-VLAN routing is centralized routing, and illustrated as below, 



    Explanation:

    • A dedicated router - up506/cenRtr is used to route the traffic between vlans;
    • All gateway SVIs on the cenRtr
    • Advantages:
      • Low resource requirement on VTEPs, which only need to know how to reach gateway. So fewer MAC and no ARP
      • Easy managed. 
    • Disadvantages:
      • Sub-optimal traffic flow. 
      • Single point failure
    Control Plane Check-up:

    1. IMET:

    On centralized router, under vlan 631, only 2 VTEPs - local and VTEP1

    up506-CentRtr#show bgp evpn route-type imet vni 631
              Network                Next Hop              Metric  LocPref Weight  Path
     * >     RD: 160.255.255.10:630 imet 631 160.255.255.10
                                    160.255.255.10        -       100     0       i Or-ID: 160.255.255.10 C-LST: 180.255.255.1
     * >     RD: 160.255.255.20:630 imet 631 160.255.255.100
                                    -                     -       -       0       i

    Similarly, under vlan 632, only 2 VTEPs - local and VTEP2

    up506-CentRtr#show bgp evpn route-type imet vni 632
              Network                Next Hop              Metric  LocPref Weight  Path
     * >     RD: 160.255.255.20:630 imet 632 160.255.255.20
                                    160.255.255.20        -       100     0       i Or-ID: 160.255.255.20 C-LST: 180.255.255.1
     * >     RD: 160.255.255.20:630 imet 632 160.255.255.100
                                    -                     -       -       0       i

    2. MAC-IP:

    up506-CentRtr#show bgp evpn route-type mac-ip vni 631
              Network                Next Hop              Metric  LocPref Weight  Path
     * >     RD: 160.255.255.10:630 mac-ip 631 444c.a8a5.1140
                                    160.255.255.10        -       100     0       i Or-ID: 160.255.255.10 C-LST: 180.255.255.1

    up506-CentRtr#show bgp evpn route-type mac-ip vni 632
              Network                Next Hop              Metric  LocPref Weight  Path
     * >     RD: 160.255.255.20:630 mac-ip 632 444c.a8a5.1141
                                    160.255.255.20        -       100     0       i Or-ID: 160.255.255.20 C-LST: 180.255.255.1

    Ping check-up:

    Host1#ping vrf EvpnHost1 160.63.2.202
    PING 160.63.2.202 (160.63.2.202) 72(100) bytes of data.
    80 bytes from 160.63.2.202: icmp_seq=1 ttl=63 time=0.157 ms
    80 bytes from 160.63.2.202: icmp_seq=2 ttl=63 time=0.112 ms
    80 bytes from 160.63.2.202: icmp_seq=3 ttl=63 time=0.144 ms
    80 bytes from 160.63.2.202: icmp_seq=4 ttl=63 time=0.106 ms
    80 bytes from 160.63.2.202: icmp_seq=5 ttl=63 time=0.133 ms

    --- 160.63.2.202 ping statistics ---
    5 packets transmitted, 5 received, 0% packet loss, time 0ms
    rtt min/avg/max/mdev = 0.106/0.130/0.157/0.021 ms, ipg/ewma 0.156/0.143 ms

    5/21/2020

    Arista EVPN VXLAN Configuration Example (2c) - Single-homing, L2 EVPN, Vlan-aware vs Vlan-based

    According to draft-krattiger-evpn-modes-interop-0, the vlan-based should interop with vlan-aware bundled MAC VRF. But I did a quick test on 4.24.0F EOS, it doesn't work obviously



    And the problem is that, flood set is not correct.

    snp261#sh l2rib input bgp floodset
    L2 RIB EVPN Input flood set:
       Vlan              Address       Type            Destination
    ---------- -------------------- ---------- -------------------
        601       0000.0000.0000        All    VTEP 160.255.255.20
        602       0000.0000.0000        All    VTEP 160.255.255.20
        611       0000.0000.0000        All    VTEP 160.255.255.20
        612       0000.0000.0000        All    VTEP 160.255.255.20
    <<<< there is no VTEP flood set for VLAN 621 and 622

    Even the "show bgp evpn route-type imet <prefix>" shows correct RT values, the but import doesn't work here. 

    wa464#sh bgp evpn route-type imet rd 160.255.255.10:621 detail
    BGP routing table information for VRF default
    Router identifier 160.255.255.20, local AS number 65162
    BGP routing table entry for imet 160.255.255.10, Route Distinguisher: 160.255.255.10:621
     Paths: 1 available
      Local
        160.255.255.10 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.10, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:620 TunnelEncap:tunnelTypeVxlan
          VNI: 621
          PMSI Tunnel: Ingress Replication, MPLS Label: 621, Leaf Information Required: false, Tunnel ID: 160.255.255.10

    Arista EVPN VXLAN Configuration Example (2b) - Single-homing, L2 EVPN, Vlan-aware

    The VLAN-based MAC VRF has RD/RT values per VLAN/VNI. But most of the time, one tenant customer uses multiple VLANs. In this case, we can use just 1 RD/RT to mark the EVPN routes, which is called VLAN-aware bundle MAC VRF, and is illustrated as below:



    Explanations:
    • Configuration is much like the VLAN-based MAC VRF
    • Configure VLAN-aware MAC-VRF under router BGP with RD/RT and it can have multiple VLANs
    • "redistribute learned" is to advertised the learnt MAC under VLAN as type-2 routes to remote EVPN peers.
    • Under interface Vxlan 1, configure VNI values for above VLANs

    Control Plane Check-up:

    1. IMET, almost same as VLAN-based, but under 1 RD/RT with 2 VNIs/VTEP

    snp261-eVtep1#sh bgp evpn route-type imet rd 160.255.255.20:610 detail
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for imet 611 160.255.255.20, Route Distinguisher: 160.255.255.20:610
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:610 TunnelEncap:tunnelTypeVxlan
          VNI: 611
          PMSI Tunnel: Ingress Replication, MPLS Label: 611, Leaf Information Required: false, Tunnel ID: 160.255.255.20
    BGP routing table entry for imet 612 160.255.255.20, Route Distinguisher: 160.255.255.20:610
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:610 TunnelEncap:tunnelTypeVxlan
          VNI: 612
          PMSI Tunnel: Ingress Replication, MPLS Label: 612, Leaf Information Required: false, Tunnel ID: 160.255.255.20

    2. MAC-IP, under 1 RD/RT, but 2 VNI for 2 VLANs

    snp261-eVtep1#sh bgp evpn route-type mac-ip rd 160.255.255.20:610 detail
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for mac-ip 611 444c.a8a5.1141, Route Distinguisher: 160.255.255.20:610
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:610 TunnelEncap:tunnelTypeVxlan
          VNI: 611 ESI: 0000:0000:0000:0000:0000
    BGP routing table entry for mac-ip 612 444c.a8a5.1141, Route Distinguisher: 160.255.255.20:610
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:610 TunnelEncap:tunnelTypeVxlan
          VNI: 612 ESI: 0000:0000:0000:0000:0000

    Ping check-up:

    Host1#ping vrf EvpnHost1 160.61.1.201
    PING 160.61.1.201 (160.61.1.201) 72(100) bytes of data.
    80 bytes from 160.61.1.201: icmp_seq=1 ttl=64 time=0.790 ms
    80 bytes from 160.61.1.201: icmp_seq=2 ttl=64 time=0.134 ms
    80 bytes from 160.61.1.201: icmp_seq=3 ttl=64 time=0.102 ms
    80 bytes from 160.61.1.201: icmp_seq=4 ttl=64 time=0.107 ms
    80 bytes from 160.61.1.201: icmp_seq=5 ttl=64 time=0.117 ms

    --- 160.61.1.201 ping statistics ---
    5 packets transmitted, 5 received, 0% packet loss, time 4ms
    rtt min/avg/max/mdev = 0.102/0.250/0.790/0.270 ms, ipg/ewma 1.000/0.510 ms

    Host2#ping vrf EvpnHost1 160.61.2.201
    PING 160.61.2.201 (160.61.2.201) 72(100) bytes of data.
    80 bytes from 160.61.2.201: icmp_seq=1 ttl=64 time=0.884 ms
    80 bytes from 160.61.2.201: icmp_seq=2 ttl=64 time=0.105 ms
    80 bytes from 160.61.2.201: icmp_seq=3 ttl=64 time=0.106 ms
    80 bytes from 160.61.2.201: icmp_seq=4 ttl=64 time=0.097 ms
    80 bytes from 160.61.2.201: icmp_seq=5 ttl=64 time=0.096 ms

    --- 160.61.2.201 ping statistics ---
    5 packets transmitted, 5 received, 0% packet loss, time 4ms
    rtt min/avg/max/mdev = 0.096/0.257/0.884/0.313 ms, ipg/ewma 1.001/0.559 ms

    Arista EVPN VXLAN Configuration Example (2a) - Single-homing, L2 EVPN, Vlan-based

    As of May 2020, the latest EOS 4.24.0F supports 2 types of L2 EVPN MAC-VRF:
    • Vlan-based
      • 1 VRF (1 RD): 1 VLAN
      • ESI = 0
    • Vlan-aware Bundled
      • 1 VRF (1 RD): n VLANs
      • ESI = VNI
    Here is the configuration of Vlan-Based L2 EVPN

    Explanations:
    •  Configure VLAN aka MAC-VRF under router BGP with RD/RT
    • "redistribute learned" is to advertised the learnt MAC under VLAN as type-2 routes to remote EVPN peers.
    • Under interface Vxlan 1, configure VNI values for above VLANs
    Control Plane Checkup:

    EVPN uses 2 types routes for L2EVPN, 1) type-3 IMET, 2) type-2 MAC-IP

    1) IMET

    snp261-eVtep1#show bgp evpn route-type imet vni 601 detail
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for imet 160.255.255.10, Route Distinguisher: 160.255.255.10:601
     Paths: 1 available
      Local
        - from - (0.0.0.0)
          Origin IGP, metric -, localpref -, weight 0, valid, local, best
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan
          VNI: 601
          PMSI Tunnel: Ingress Replication, MPLS Label: 601, Leaf Information Required: false, Tunnel ID: 160.255.255.10
    BGP routing table entry for imet 160.255.255.20, Route Distinguisher: 160.255.255.20:601
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan  << RT to control import
          VNI: 601  << VNI for this VLAN
          PMSI Tunnel: Ingress Replication, MPLS Label: 601, Leaf Information Required: false, Tunnel ID: 160.255.255.20 << VTEP ID

    2) Flood-set, the above IMET prefix is used to form the flood-set

    snp261-eV1.16:27:32#show vxlan flood vtep vlan 601
              VXLAN Flood VTEP Table
    --------------------------------------------------------------------------------
    VLANS                            Ip Address
    -----------------------------   ------------------------------------------------
    601                             160.255.255.20

    snp261-eV1.16:23:04#show l2rib output floodset vlan 601
    L2 RIB Output flood set:
    Source: Local Dynamic, Local Static, BGP, VXLAN Static, VXLAN Dynamic
       Vlan              Address       Type            Destination
    ---------- -------------------- ---------- -------------------
        601       0000.0000.0000        All    VTEP 160.255.255.20

    snp261-eV1.16:27:02#show l2rib input bgp floodset vlan 601
    L2 RIB EVPN Input flood set:
       Vlan              Address       Type            Destination
    ---------- -------------------- ---------- -------------------
        601       0000.0000.0000        All    VTEP 160.255.255.20

    3) Type-2 MAC-IP EVPN Route

    snp261-eVtep1#show bgp evpn route-type mac-ip vni 601 detail
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for mac-ip 444c.a8a5.1140, Route Distinguisher: 160.255.255.10:601
     Paths: 1 available
      Local
        - from - (0.0.0.0)
          Origin IGP, metric -, localpref -, weight 0, valid, local, best
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan
          VNI: 601 ESI: 0000:0000:0000:0000:0000
    BGP routing table entry for mac-ip 444c.a8a5.1141, Route Distinguisher: 160.255.255.20:601  << mac address
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan << RT to control import
          VNI: 601 ESI: 0000:0000:0000:0000:0000 << VNI

    4) MAC table vs EVPN prefixes

    snp261-eV1.16:29:56#show mac address-table interface vxlan 1 vlan 601
              Mac Address Table
    ------------------------------------------------------------------

    Vlan    Mac Address       Type        Ports      Moves   Last Move
    ----    -----------       ----        -----      -----   ---------
     601    444c.a8a5.1141    DYNAMIC     Vx1        1       4:53:10 ago

    5) Clear MAC on remote VTEP to simulate MAC aging out

    wa464-eVtep2#clear mac address-table dynamic vlan 601 << clear MAC

    snp261-eVtep1#show bgp evpn route-type mac-ip vni 601 detail << NO evpn type-2 prefix
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for mac-ip 444c.a8a5.1140, Route Distinguisher: 160.255.255.10:601
     Paths: 1 available
      Local
        - from - (0.0.0.0)
          Origin IGP, metric -, localpref -, weight 0, valid, local, best
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan
          VNI: 601 ESI: 0000:0000:0000:0000:0000

    snp261-eVtep1#show mac address-table interface vxlan 1 vlan 601 << no MAC entry
              Mac Address Table
    ------------------------------------------------------------------

    Vlan    Mac Address       Type        Ports      Moves   Last Move
    ----    -----------       ----        -----      -----   ---------

    Data Plane Checkup: 

    host1 under VTEP1 ping host2 behind VTEP2

    Host1#ping vrf EvpnHost1 160.60.1.201
    PING 160.60.1.201 (160.60.1.201) 72(100) bytes of data.
    80 bytes from 160.60.1.201: icmp_seq=1 ttl=64 time=0.135 ms
    80 bytes from 160.60.1.201: icmp_seq=2 ttl=64 time=0.100 ms
    80 bytes from 160.60.1.201: icmp_seq=3 ttl=64 time=0.092 ms
    80 bytes from 160.60.1.201: icmp_seq=4 ttl=64 time=0.088 ms
    80 bytes from 160.60.1.201: icmp_seq=5 ttl=64 time=0.089 ms

    --- 160.60.1.201 ping statistics ---
    5 packets transmitted, 5 received, 0% packet loss, time 0ms
    rtt min/avg/max/mdev = 0.088/0.100/0.135/0.021 ms, ipg/ewma 0.128/0.117 ms


    VLAN-based: RD/RT vs VNI = 1:1

    From the below output, the different VLANs have different RD and RT values, so 1:1 relationship. (In our case, only one host simulates multiple hosts under different VLANs).

    snp261-eV1.18:18:45#sh bgp evpn route-type mac-ip vni 601 detail
    BGP routing table entry for mac-ip 444c.a8a5.1141, Route Distinguisher: 160.255.255.20:601
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:601 TunnelEncap:tunnelTypeVxlan
          VNI: 601 ESI: 0000:0000:0000:0000:0000

    snp261-eV1.18:18:49#sh bgp evpn route-type mac-ip vni 602 detail
    BGP routing table entry for mac-ip 444c.a8a5.1141, Route Distinguisher: 160.255.255.20:602
     Paths: 1 available
      Local
        160.255.255.20 from 160.255.255.1 (180.255.255.1)
          Origin IGP, metric -, localpref 100, weight 0, valid, internal, best
          Originator: 160.255.255.20, Cluster list: 180.255.255.1
          Extended Community: Route-Target-AS:65100:602 TunnelEncap:tunnelTypeVxlan
          VNI: 602 ESI: 0000:0000:0000:0000:0000

    5/20/2020

    Arista EVPN VXLAN Configuration Example (1c) - eBGP Overlay

    After the eBGP Undelay is up, we can move on to the eBGP Overlay.

    Topology and configuration:


    Explanations:
    • The BGP AS# of spine and leaf routes are different, so need to use BGP local-as feature to form iBGP peering
    • The benefit of BGP RR is that, the NH of EVPN updates are unchanged
    • Also need to disable the RR under address family ipv4 unicast or configure "no bgp default ipv4-unicast" under router bgp, otherwise you will see the following failed neighbor
    bn303-eSpine1#show ip bgp summary
    BGP summary information for VRF default
    Router identifier 180.255.255.1, local AS number 65100
    Neighbor Status Codes: m - Under maintenance
      Description              Neighbor         V  AS           MsgRcvd   MsgSent  InQ OutQ  Up/Down State   PfxRcd PfxAcc
    ....
      eOverL-V1                160.255.255.10   4  65100            552       545    0    0 07:40:16 Estab(NotNegotiated)
      eOverL-V2                160.255.255.20   4  65100            555       543    0    0 07:40:16 Estab(NotNegotiated)

    And in the following blogs of various L2/L3 vxlan evpn setup, we don't need to touch spine routers anymore
    • From control plane point of view, spine1 only reflects BGP EVPN updates among leaf routers w/o VXLAN interface or VRF, which means it doesn't need to understand or import the content. 
    • From data plane point of view, spine1 only forwards the IPv4/VXLAN packets by the source/destination address are leafs' loopback address. 
    Verifications:

    1) BGP EVPN peering to VTEP1/2 are up

    bn303-eSpine1#show bgp evpn summary
    BGP summary information for VRF default
    Router identifier 180.255.255.1, local AS number 65100
    Neighbor Status Codes: m - Under maintenance
      Description              Neighbor         V  AS           MsgRcvd   MsgSent  InQ OutQ  Up/Down State   PfxRcd PfxAcc
      eOverL-V1                160.255.255.10   4  65100            581       573    0    0 00:20:21 Estab   2      2
      eOverL-V2                160.255.255.20   4  65100            585       571    0    0 00:20:21 Estab   2      2

    Arista EVPN VXLAN Configuration Example (1b) - eBGP Underlay

    Before start, let's talk a bit regarding the underlay vs overlay. This topic is very well covered in the EVPN Deployment Guide.  Here is my easy understanding:
    • The underlay is:
      • EBGP between Spine and Leafs over P2P links. Of course, IGP is an option. 
      • To advertise the routing information of the VTEPs' and Spines' loopbacks for
        • EVPN peering sessions
        • The source/destination address Vxlan data traffic
      • You can use different loopback for EVPN and Vxlan
        • In EVPN Deployment Guide (page 20), lo0 for EVPN peering, lo1 for VxLAN
        • In this blog, I use the same loopback160 for both
    • Overlay control plane is BGP EVPN, to 
      • Register the VTEP (type 3), it is like L2vpn autodiscovery. 
      • Carry EVPN prefixes L2 and L3
    Topology and Configuration:



    Explanation:
    • EVPN is only supported in BGP multi-agent mode. 
    • Simple EBGP peering over b2b ethernet interfaces
    • All routers advertise the loopback160 /32 address
    • Please note, the spine routers/RR have a route-map to control, so that only /32 loopback prefixes within 160.255.255.0/24 are advertised. 
      • The b2b interface addresses are not needed in the data plane, so unnecessary for routing protocol. 
      • And this eBGP setup implies that interface address can be duplicated in different POP/DC. A big plus for automation 
    Control Plane Verification:
    1) show ip bgp sum on VTEP1, and please note only 2 bgp routes

    snp261-eVtep1#sh ip bgp sum
    BGP summary information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    Neighbor Status Codes: m - Under maintenance
      Description              Neighbor         V  AS           MsgRcvd   MsgSent  InQ OutQ  Up/Down State   PfxRcd PfxAcc
      eUnder-Sp1               160.1.10.1       4  65100          23840     23848    0    0   13d17h Estab   2      2

    2) show ip bgp on VTEP1 to check the bgp prefixes, 1 for spine, 1 from VTEP2

    snp261-eVtep1#show ip bgp
              Network                Next Hop              Metric  LocPref Weight  Path
     * >      160.255.255.1/32       160.1.10.1            0       100     0       65100 i
     * >      160.255.255.10/32      -                     -       -       0       i
     * >      160.255.255.20/32      160.1.10.1            0       100     0       65100 65162 i

    3) show bgp prefix of VTEP2's loopback on VTEP1

    snp261-eVtep1#sh ip bgp 160.255.255.20
    BGP routing table information for VRF default
    Router identifier 160.255.255.10, local AS number 65161
    BGP routing table entry for 160.255.255.20/32
     Paths: 1 available
      65100 65162
        160.1.10.1 from 160.1.10.1 (180.255.255.1)
          Origin IGP, metric 0, localpref 100, weight 0, received 23:22:11 ago, valid, external, best
          Rx SAFI: Unicast

    4) show ip route on VTEP1 to ensure bgp route in routing table

    snp261-eVtep1#show ip route 160.255.255.20/32
     B E      160.255.255.20/32 [200/0] via 160.1.10.1, Ethernet13

    5) VTEP1 pings VTEP2's lo160 - 160.255.255.20

    snp261-eVtep1#ping 160.255.255.20 source lo160
    PING 160.255.255.20 (160.255.255.20) from 160.255.255.10 : 72(100) bytes of data.
    80 bytes from 160.255.255.20: icmp_seq=1 ttl=63 time=0.248 ms
    80 bytes from 160.255.255.20: icmp_seq=2 ttl=63 time=0.128 ms
    80 bytes from 160.255.255.20: icmp_seq=3 ttl=63 time=0.087 ms
    80 bytes from 160.255.255.20: icmp_seq=4 ttl=63 time=0.082 ms
    80 bytes from 160.255.255.20: icmp_seq=5 ttl=63 time=0.091 ms

    --- 160.255.255.20 ping statistics ---
    5 packets transmitted, 5 received, 0% packet loss, time 0ms
    rtt min/avg/max/mdev = 0.082/0.127/0.248/0.062 ms, ipg/ewma 0.179/0.184 ms

    At this step, we are pretty sure the underlay is ready because the overlay bgp evpn peering is based on VTEPs' loopback interfaces.

    Arista EVPN VXLAN Configuration Example (1a) - Overview

    I am starting a series of blogs on the Arista EVPN Vxlan configuration and will cover the following topics:
    • Underlay/Overlay BGP Configuration (No IGP involved)
    • L2 EVPN vs L3
      • L2 VLAN-based vs L2 VLAN-aware
      • Symmetric and asymmetric routing
    • Single-homing and multi-homing
      • Multi-homing: MLAG vs EVPN active/active
    • 2 major Arista platforms: Jericho and Trident families
    • Router reflector and router server
    • Inter-VPN solution
    • EVPN inter VRF route leak
    • Some other advanced features
      • Dynamic BGP peering
    Pre-requisites: 2 things need to be taken care of before EVPN configuration:
    • BGP multi-agent mode, EVPN is ONLY supported with this mode
      • service routing protocols model multi-agent
      • And need to reboot device to make this effective
    • Hardware setup:
      • For Trident 2 (7050*X), Tomahawk (7060*X) family devices like , recirculation must be enabled
      • For Arad (7280E, 7500E)/Jericho (7280R*, 7500R*) series device, select vxlan-routing TCAM profile
    List of reference:
    What's the difference between this series of blog and the above official documents?
    • Configuration and trouble-shooting focused, no theory. The above links did a good job on the theory explanation, so I don't need to waste time and effort here. 
    • Simplest topology and step by step configuration
      • Starting with 2 single-homing VTEPs + 1 Spine + RR
      • L2 EVPN (vlan-based, vlan-bundle-aware) and L3 EVPN
      • 2 dual-homing VTEPs by MLAG or EVPN A/A
      • IPv4/v6 overlay
      • Adding 1 Spine for ECMP and RS
      • Only necessary and user-input configuration (no BGP max routes or link speed)
    • Hardware and scale information and consideration
    • Convergence/switchover time