2012-02-08

Nearly there

Well, progress again today.

We now have all our transit and peering point connections on the new rack.

We have ADSL customers connected via the new rack, and platform RADIUS to carriers from it all working.

We have the RADIUS servers and DNS servers and syslog servers in there too.

We have our new Ethernet hub in there, but not yet tested a live connection.

The main thing left now is to get the direct connections to other ISPs moved, and we have the links in place ready to do that. The delay there is co-ordinating with each interconnect separately and moving at a convenient time with little or no disruption to their services.

Once the old rack is cleared out, we have to set up the other host link in the new rack (moving from old rack), and that will be it!

All good fun.

TR-069 ACS

Apart from all the upgrades we are doing, I do have other work on as well, and one of the projects that is quite urgent is making a TR-069 server.

I am reading through the spec now - and it is going to be fun.

Of course, there are existing servers, well, at least one, open source, but I do like to re-invert the wheel if I can.

The main thing is that it will be completely in C so should be quite efficient.

Why? Well, for a start, we have these TG582n technicolor routers, and they (a) support TR-069, and (b) have a fairly crap web interface for config. So the plan is to make our control pages for DSL have a whole router config page that allows customers to config their routers centrally and it update the router via TR-069.

Apart from allowing us to make a web config page for the router, and perhaps other routers, it will also allow for replacement routers to have the exact same config, and even for routers to get the config if factory reset.

Obviously we are all about giving our end users choice, and this will be optional. However, even for our more techie customers, this is likely to appeal I expect.

So, the question is, do I make this TR-069 server open source? I may. It does mean writing it slightly differently. If it was only an in-house tool it would be integrated in to our databases and systems very tightly. But it may be more fun to make it a general purpose tool.

If any of you are interested in this, do let me know, and the main features you are looking for.

We are looking for (a) ability to send a new config to the router (b) ability to upgrade firmware on the router, and that is pretty much it!

2012-02-07

There's always one!

Lots of progress today.

We have LONAP (another peering point) connected on to the new rack.

We have BE wholesale connected as well. It did mean the FireBrick team working on changes in the TCP stack, BGP code, config and documentation to add a non standard forced TTL on BGP peers, but that done - it is working.

We should be able to move the direct connection customers from tomorrow.

Of course we found a fun one - with sessions spread over two LNSs there was, of course, a customer that did not work. It turns out they were the only remaining customer using an old feature called closed user group which allows restricted access between sites. It only works where the sites are on the same LNS else no connectivity at all. And they were not on the same LNS. It is simple to fix, and just means using a prefix on their user name, so sorted. But there had to be one exception, and I knew it would either be one of Mike's customers or one of Kev's. It was one of Mike's... Well done Mike. It is funny how you can guess who the unusual configuration lines are with :-)

Anyway, we are waiting on BT to put an ISDN line in for the old host link moving to the new rack (to monitor the fibres). We have one transit feed still to move, but that will be this week. There is a good chance that in a few days time we will have nothing actually using the old rack at all.

Then we'll turn the old rack off at least a day before we start pulling kit out, just to see who screams :-)

Getting there.

Well done team.

2012-02-06

Silly set backs today

Well, lots of silly things.

One customer is logging in using upper case on one line and lower case on the other so split between two LNSs as the system uses a hash of the username to decide. Hence my script to check if we have any sites plit over two LNSs wanted to kill their lines every 60 seconds. It is the only customer and they are changing their login. D'Oh!

We have our favorite telco that realise they did not put enough copper pairs to our rack! They need them for monitoring the fibres, but one lot is monitored using ADSL and one is monitored using ISDN. That will delay moving the old host links to the new rack.

We are getting our management LAN DSL installed, with another ISP, as you do. I won't say who, but how quaint: Paper order form. Paper DD form. So waiting for my signature in the office now. No IPv6. And this is not a small ISP...

We think we have sorted PI space customers now, but may need some more sophisticated changes to the source filtering on L2TP to manage it "neatly".

Of course, as blogged, arguments with one supplier wanting non standard BGP links. They are making an exception for now - thanks!

Sounds like most of the remaining jobs should be sorted in next few days, but nice to let things settle a bit in the mean time.

Behind schedule, as always, but we allowed lots of slack just in case.

Oh, and we are putting more bandwidth on 21CN. Lots of migrates from 20CN this month. By more I mean shit loads more... I am really trying to get my stats up to the full 100.0% no dropped packets, if I can.

So, a fun week to follow.

Non standard BGP

We have an interesting case with one of our carriers. Looks like we have worked around it for now, but it is rather odd as they are requiring non standard BGP TCP/IP in links to them.

They are requiring us to send all the BGP TCP packets with a TTL of 1

What is interesting is that some of the big routers do indeed do this, but doing so is against the recommendations for TCP/IP which recommends a TTL of 64. BGP itself makes no mention of TTL. There are Internet standards that say the TTL must be at least the Internet diameter even. So naturally, our BGP does in fact follow this standard and uses a TTL of 64.

Of course, using TTL of 1 was a silly thing anyway as anyone could spoof BGP with a TTL of 1 by setting it to a suitable higher value, though if the reply was TTL of 1 they do not get far. The issue that came up is anyone can spoof a convincing TCP RST with a TTL of one and shut down BGP sessions. This problem is now recognised in other Internet standards which document TTL security where one sets a TTL of 255 and the far end checks it is still 255. Remotely spoofing a TTL of 255 is impossible without compromising the local routers somehow. So that works. Indeed our routers support TTL security.

It seems this is some fire-walling rule, and to be honest I have never seen anyone fire-walling based on TTL. It is not clear what they are trying to protect against. They allow L2TP with normal TTLs.

A simpler firewall would be to do the same as other carriers and have access lists covering which of our IPs can talk to which of their IPs, and not fire-walling on TTL.

This is the same bunch that allowed MAC spoofing on PPPoE links to disrupt and even monitor other people's DSL lines. Thankfully they are fixing that.

It seems however they are insisting that any future services we buy must use this non standard BGP to connect to them.

It is very brave of them.

I guess following the standards is a key factor in deciding which suppliers we will use for new services.

P.S. FireBrick have, of course, modified the TCP stack and BGP and config on the FireBrick BGP routers to support this non standard mode of operation as well as standard TTL security.

Breaking new ground

There are a couple of completely new things that are being used in anger this week. We have done various tests last week, but this is for real now.
Firstly we are now running a dual live LNS system using two of the new LNSs. We have customers spread over the two LNSs based on their login. This means, in theory, bonded customers are on the same LNS. If not, then you get working service but not the bonded download. We are working on ways to pick up any that go wrong, and we think some BE lines ended up on the wrong LNS. A simple ppp-kill will move to the right one.

This is also the first time we have run the LNSs in the route-reflector so they see all other routes. They used to announce connections and use the core routers as a gateway - now they see all routes and can send via the right external gateways for outgoing traffic.

Both of these could have unexpected side effects. We have seen some with customers that have PI space (their own IP blocks).

So anyone with something odd, please let tech staff know and we can resolve the problems as they come up. We managed to sort one PI issue within minutes of being reported on a Sunday night.

In the mean time we'll get on moving the other transit, peering points and direct links over during this week.

2012-02-05

Upgrade progress (LINX and transit)

We have managed to move one transit and our LINX peering over now. We have all customers moved over on to the new LNSs, except the wholesale ones.

We even found why data SIMs were not showing graphs and sorted.

So has been a fun weekend.

Next week we get other transit feeds, other peering points, and a whole load of direct peering - which is going to take some co-ordinating.

The major jobs are sorted though, and all is looking very good.

You do then hit fiddly things like making sure nagios is watching the right boxes, and ensuring your cacti graphs are all running on the new boxes, and checking all the management LAN works, and the backup out-of-band access works, and the administration passwords are all set correctly with the right access lists. For the most part it is copy and paste, but you have to then test everything carefully just in case. A never ending set of silly little details.

At some point we want to go in there on a Sunday and check the dual power, which should be seamless. We also want to check that taking out a whole side of the network (turning off a switch) recovers. That will take some lines out for a few minutes we expect. We need to make a list of carefully defined tests and make sure people know we are buggering about.

Ideally, at some point, we should test turning off the whole power, and then back on, and seeing how quickly everything recovers. I am not sure if we will do that or not - it is a bit disruptive.

But if you don't test the contingencies they bite you when something does break.

We'll post details of what tests are being done when.

Clocks

Some time geeks (should I say Time Lords) checked out my clocks. Seems they are impressed, saying sub microsecond. I have spent all day tryi...